Tag: Reduce Error Rate in the Back Office

  • Cutting Contract Review Errors by 60% in a Two-Week B2B SaaS Pilot

    1. Baseline Error Rate Is the Real KPI

    The finance team at a 2,000+ employee B2B SaaS company in Vienna processes roughly 1,200 contracts per month. Each one passes through a manual review queue where an analyst extracts termination clauses, liability caps, and auto-renewal flags into the ERP. The baseline error rate sits at 5.2%: a missed auto-renewal date or a misread liability cap ends up in the system and surfaces three months later during a renewal dispute. A two-week pilot with a dedicated AI team replaced the manual extraction step with a LangGraph pipeline that parses PDFs, extracts 14 structured fields, and writes the result to a staging table via a custom REST API. The measured error rate dropped to 1.8% on the pilot’s 300-contract sample, and cycle time per contract fell from 11 minutes to 90 seconds of model time plus 4 minutes of human approval. The pilot did not touch the production ERP; it ran on a read-only copy of the contract repository and output to a sandbox workspace in the CRM.

    2. LangGraph Handles the Multi-Step Extraction

    The extraction pipeline runs on LangGraph, not a single LLM call. The graph has five nodes: PDF ingestion (PyMuPDF for text-layer PDFs, Tesseract OCR fallback for scanned documents), clause segmentation (a fine-tuned classifier that splits the document into 8–12 logical sections), field extraction (GPT-4o for high-accuracy fields like liability caps, Llama 3 70B on the client’s own GPU for fields containing personal data), confidence scoring, and output formatting. The REST API endpoint POST /v1/extract accepts a multipart PDF upload and returns a JSON object with 14 fields, each carrying a confidence score between 0 and 1. Fields below 0.90 route to a human reviewer in the existing helpdesk queue; fields at or above 0.90 auto-populate the staging table. Webhooks fire on completion so the finance team’s dashboard updates without polling. The entire pipeline runs on the client’s AWS eu-central-1 region, keeping data within Austria’s borders.

    3. Two Weeks Is Enough for a Measured Pilot

    The pilot ran for exactly 14 calendar days. Days 1–3: process audit. The AI team shadowed three finance analysts, logged every manual step, and identified the 14 fields that caused the most downstream errors. Days 4–6: data preparation. The team pulled 300 historical contracts from the repository, had two analysts independently annotate the 14 fields, and resolved disagreements to build a gold-standard test set. Days 7–10: pipeline build and tuning. The LangGraph workflow was assembled, the extraction prompt was iterated four times, and the confidence threshold was calibrated so that the false-negative rate (a wrong value auto-approved) stayed below 0.5%. Days 11–14: measurement. The pipeline ran on the 300-contract set, and the team compared field-level accuracy against the gold set, measured cycle time, and produced a before/after report. The report included a cost model: at 1,200 contracts per month, the pilot’s error reduction translated to an estimated EUR 18,400 in avoided dispute costs per quarter.

    4. Human-in-the-Loop Is Non-Negotiable

    The model does not replace the analyst; it removes the 11 minutes of copy-paste and field-mapping that precede the actual judgment call. The human-in-the-loop design is explicit: the model drafts the 14 extracted fields, the analyst reviews them in a purpose-built UI that highlights low-confidence fields in amber, and the analyst approves or corrects before the record writes to the ERP. For a B2B SaaS company, the highest-risk fields are termination notice periods and liability caps, because a wrong value here has direct financial consequences. The pilot’s measurement showed that 78% of fields required no human correction, 19% needed a single-field edit, and 3% required a full re-extraction. The analyst’s role shifted from data entry to exception handling, which freed roughly 6.5 hours per analyst per week. The dedicated AI team operated the pipeline during the pilot, monitored confidence drift, and tuned the prompt when a new contract template appeared in the sample.

    5. The Integration Is a Thin REST Layer

    The pilot’s REST API and webhook architecture was designed to plug into the client’s existing stack without replacing it. The extraction service exposes a stateless POST /v1/extract endpoint that the finance team’s internal tool calls via a simple HTTP request. On completion, a webhook POSTs the result to the client’s CRM (Salesforce) and ERP (SAP S/4HANA) through their respective API endpoints. No middleware, no new database, no replacement of the existing document management system. The client’s IT team reviewed the API contract in day 2 of the pilot and approved the integration scope. The model-agnostic design meant the team could swap GPT-4o for Llama 3 on the client’s GPU for any field that contained personal data, without changing the API contract or the downstream integration. This matters for a 2,000+ employee firm where IT governance requires that no new SaaS dependency is introduced for a pilot that may not scale.

    6. What the Pilot Does Not Cover

    The pilot’s 1.8% error rate is not the end state. The team’s rollout plan, presented in the final pilot report, targets a 0.9% error rate within 90 days of production deployment. The path: expand the gold-standard test set from 300 to 2,000 contracts, add a second extraction pass for fields with confidence between 0.80 and 0.90, and introduce a feedback loop where analyst corrections are logged and used to fine-tune the clause-segmentation classifier. The dedicated AI team continues to operate the pipeline in production, monitoring a dashboard that tracks field-level accuracy, confidence distribution, and cycle time per contract. The B2B SaaS firm’s finance director approved the rollout on the basis of the pilot’s measured numbers, not a projection. The two-week window was sufficient because the scope was narrow: one document type, 14 fields, one team, one measurement. Expanding to multi-party agreements or adding a second document type (e.g., purchase orders) would require a second pilot of similar duration.

  • Cutting Back-Office Error Rates 47% in a 24-Person Austrian E-Commerce Firm

    Background: A 24-Person E-Commerce Operator in Vienna

    This case study is a composite built from patterns Forfis has observed across multiple e-commerce and retail engagements in Tier-1 European markets. No named customer is represented. The company, the metrics, and the timeline are drawn from recurring patterns in the field, not from a single identifiable client.

    The company is a 24-person e-commerce operator based in Vienna, selling home goods and small appliances across Austria and Germany. It runs a Shopify storefront, a NetSuite ERP, and a Zendesk helpdesk. The back-office team of six handles invoice processing, order data entry, and first-line support triage. The company holds ISO 27001 certification, a requirement for its B2B wholesale channel. The CTO is a former infrastructure engineer who has run the stack for four years and is comfortable with REST APIs and webhooks but has no prior AI engineering experience. The team is in the scaling phase: revenue has grown 60% year-over-year, but the back-office error rate has climbed from 3.2% to 7.8% because the same six people are processing 40% more volume without additional headcount.

    Challenge: Error Rates Climbing, Headcount Flat, ISO 27001 in the Way

    The trigger was a quarterly audit that flagged a 7.8% error rate in invoice and order data entry, up from 3.2% eighteen months earlier. Each error required a manual correction, an average of 14 minutes of back-office time, and in 12% of cases triggered a customer-facing refund or credit. The support team was also drowning: 340 tickets per week, 68% of which were first-response queries that a knowledge base search could have resolved without a human. The CTO had two constraints. First, ISO 27001 required that no customer PII or payment data leave the company’s infrastructure without a documented data-processing agreement. Second, the board had set a 12-week deadline to show measurable improvement before the next funding round. The CTO needed a fixed-scope engagement, not an open-ended consulting retainer. The scope had to cover three things: reduce the back-office error rate, cut first-response time on support tickets, and give the team a searchable internal knowledge base over their own documentation and CRM records.

    Approach: A 12-Week Integration Sprint on LangChain and LangGraph

    Forfis ran a two-week process audit across the back-office and support functions. The audit identified three workflows worth automating: invoice data extraction from PDF and email attachments, support ticket triage and first-response drafting, and internal knowledge search over the company’s 1,400-page product documentation and 8,200 closed support tickets. The fixed-scope pilot targeted all three, delivered as a single integration sprint over 12 weeks.

    The architecture used LangChain for prompt chaining and tool invocation, and LangGraph for the stateful, cyclic execution graphs that implement the human-in-the-loop approval pattern. The extraction pipeline ingested invoices via a custom REST API endpoint and webhooks from the email gateway. Each extracted field was scored by a predictive scoring model trained on 14 months of historical invoice data; scores below a 0.85 confidence threshold routed the document to a human reviewer. The knowledge search used retrieval-augmented generation over the company’s documentation, indexed into a vector store and updated via webhooks whenever a new document was added to the CRM. Model inference used OpenAI and Anthropic APIs for the LLM layer; the vector store and scoring model ran on the client’s own hardware to satisfy the ISO 27001 data-residency requirement. Every pipeline step logged input, output, and timestamp to an audit trail.

    Outcome: Measured Baseline Shifts in Six Weeks

    The pilot ran for six weeks after the build phase, with a two-week shadow period for the predictive scoring model before it moved to assisted mode. The measured results, compared against the pre-pilot baseline:

    • Invoice data entry error rate dropped from 7.8% to 4.1%, a 47% reduction. The remaining errors were concentrated in handwritten invoices, which the pipeline flagged for manual review rather than auto-accepting.
    • Average cycle time per invoice fell from 11.3 minutes to 6.2 minutes, a 45% reduction.
    • First-response time on support tickets dropped from 4.2 hours to 1.8 hours. The RAG-based first-response agent handled 52% of tickets without a human, with a 91% customer satisfaction score on those auto-resolved tickets.
    • Internal knowledge search reduced the time a support agent spent searching documentation from an average of 3.4 minutes per query to 0.9 minutes, a 73% reduction.
    • Back-office headcount remained at six. The team redirected the saved time to handling the 40% volume growth without hiring.

    The ISO 27001 audit trail was complete: every document processed, every model inference call, and every human approval decision was logged with a hash and timestamp. The client’s ISO 27001 certification was renewed without findings related to the new pipeline.

    Lessons for Teams Scaling AI Across Departments

    • Scope the pilot to one workflow per department, not one workflow total. The audit identified three workflows, but the pilot treated them as three parallel tracks with a shared architecture. Trying to sequence them would have blown the 12-week deadline. The shared LangGraph state machine made the parallel tracks manageable.

    • Run the predictive model in shadow mode for at least two weeks before assisted mode. The first week of shadow scoring revealed that the model’s confidence calibration was off by 0.12 on the 0.80-0.90 band. Without the shadow period, the team would have routed 18% more documents to human review than necessary, eroding the time savings.

    • Build the ISO 27001 audit trail into the pipeline from day one, not as a post-hoc compliance layer. The logging was implemented in the first week of the build, alongside the extraction logic. Retrofitting it after the pilot would have required re-running the entire pipeline on historical data, which the client did not want to do.

    • Use webhooks for the RAG index update, not a nightly batch job. The support team noticed that documents added to the CRM during the day were not searchable until the next morning. Switching to a webhook-triggered index update on document save cut the staleness window from 14 hours to under 90 seconds.

    • Keep the model layer swappable. The client asked in week 8 whether they could move the LLM inference to a self-hosted Mistral 7B model to reduce per-token costs. Because the LangChain abstraction isolated the model call, the switch was a configuration change, not a rewrite. The cost per 1,000 tokens dropped from EUR 0.03 to EUR 0.004 on the client’s existing GPU server.

  • Fixed-Scope Pilot vs. In-House Build: Lead Qualification for a UK Fintech

    What Is Being Compared

    The two options are distinct in scope and risk profile. Option A is a fixed-scope pilot delivered by an external product studio: a 6-8 week engagement on one workflow—lead qualification—using the Anthropic Claude API as the model layer, integrated via custom REST API and webhooks into the existing CRM. The studio handles technical planning, product design, and full-cycle development. The pilot ships with a measured before/after baseline on cycle time and error rate. Option B is a fully in-house build: the company’s own engineering team designs, develops, and operates the agent, using the same model API or an open-weight model on internal hardware. The in-house team owns the architecture, the integration, and the ongoing operation. Both options target the same use case—lead qualification for a 201-500 employee fintech in the UK—but they differ in who bears the delivery risk, how fast the first working system ships, and what the company must maintain after the pilot.

    Criteria for the Comparison

    The comparison is judged against seven criteria that matter to a fintech scaling operations without new hires:

    • Time to first working system — how many weeks from kickoff to a live agent handling real leads.
    • Total cost of ownership over 6 months — including model API costs, integration work, and ongoing operation.
    • PCI DSS scope impact — whether the agent’s data boundary touches cardholder data and what that means for compliance.
    • Error rate reduction — the measured delta in misclassified leads between the manual baseline and the agent.
    • Cycle time reduction — the measured delta in time from lead creation to qualified status.
    • Vendor lock-in — how easily the company can switch model providers or take the system in-house after the pilot.
    • Operational burden — who monitors, tunes, and maintains the agent after the pilot ends.

    Comparison Table

    Criterion Option A: Fixed-Scope Pilot (External Studio) Option B: In-House Build
    Time to first working system 6-8 weeks from kickoff; studio has delivery templates and prior fintech experience 12-16 weeks minimum; team must design architecture, build integration, and tune the model from scratch
    Total cost over 6 months Fixed pilot fee (typically £25,000-£40,000) plus Anthropic API usage (approx. £1,500-£3,000/month at 500-1,000 leads/month); no new hires 2-3 FTEs at £60,000-£80,000/year each plus API costs; total £150,000-£250,000 over 6 months including salaries
    PCI DSS scope impact Studio designs data boundary to exclude cardholder data; client retains compliance ownership Same design principle, but in-house team must validate the boundary against PCI DSS 4.0 requirements; no external review
    Error rate reduction Measured in pilot; studio ships with baseline and delta report; typical delta: 30-50% reduction in misclassification Measured after build; no external baseline; team must design the measurement framework themselves
    Cycle time reduction Measured in pilot; typical delta: 40-60% reduction in time-to-qualified Measured after build; no external baseline; team must design the measurement framework themselves
    Vendor lock-in Low: model-agnostic architecture; client can switch to OpenAI or an open-weight model post-pilot Low: in-house team controls the stack; no external dependency
    Operational burden Studio provides handover documentation and a 30-day post-pilot support window; client takes over operation In-house team owns all operation, monitoring, and tuning from day one

    Scenario-by-Scenario Verdict

    When Option A wins: The company has no dedicated AI engineering team and needs a working lead qualification agent within 6-8 weeks to hit a quarterly sales target. The fixed-scope pilot removes delivery risk: the studio has delivered similar systems for fintech and payments clients in Tier-1 markets, and the pilot’s measured baseline gives the sales team a concrete number to report to leadership. The 6-month timeline is tight for an in-house build, and the pilot’s fixed fee is a smaller commitment than hiring 2-3 engineers. For a 201-500 employee company where every new hire is a significant cost, the pilot’s cost profile is easier to justify.

    When Option B wins: The company already has a strong engineering team with experience in API integrations and LLM applications, and the lead qualification workflow is one of several AI initiatives the team is building. The in-house build gives the team full control over the architecture, which matters if the company plans to extend the agent to other workflows (invoice processing, document extraction) over the next 12-18 months. The in-house team can also choose to run an open-weight model on internal hardware if the data residency requirements tighten, without renegotiating a vendor contract.

    Recommendation

    For a 201-500 employee UK fintech with a 6-month timeline and no dedicated AI engineering team, Option A—the fixed-scope pilot on the Anthropic Claude API—is the better fit. The pilot’s 6-8 week delivery window fits the 6-month timeline with room for a rollout phase after the pilot. The fixed fee is a smaller financial commitment than hiring 2-3 engineers, and the studio’s prior experience with fintech and payments clients in Tier-1 markets reduces the risk of a failed pilot. The measured baseline on cycle time and error rate gives the sales team a concrete business case for scaling. The model-agnostic architecture means the company is not locked into Anthropic; if the data residency requirements change, the team can switch to an open-weight model on internal hardware without rebuilding the integration. The in-house build is the right choice only if the company already has the engineering capacity and the lead qualification agent is part of a broader AI roadmap that justifies the longer build time and higher cost.

  • AI Contract Review for Logistics: Cut Back-Office Errors by 50% in 6 Months

    1. Baseline Measurement Before You Touch a Single Clause

    Logistics firms with 201-500 employees process 500-2,000 carrier agreements, customs declarations, and service contracts monthly. Manual review by legal and compliance staff takes 15-30 minutes per document, with an 8-12% error rate on clause identification. A RAG-based contract assistant reduces this to 3-5 minutes per document with under 2% error rate. The system indexes templates and precedents from Confluence, extracts key clauses, flags deviations from standard terms, and routes exceptions to human reviewers. For a team of 12 legal staff, this saves 15-20 hours weekly, shifting focus from data entry to strategic risk assessment. The 6-month timeline includes a 4-week audit, 6-week pilot on one contract type, and 14-week rollout with measurable checkpoints at each phase.

    2. On-Premise Open-Weight Models Keep Regulated Data In-Building

    Logistics contracts often contain customs declarations, hazardous material certifications, and client NDAs with strict data residency clauses. Sending these to external APIs like OpenAI or Anthropic may violate contractual or regulatory obligations. Open-weight models like Llama 3 or Mistral deployed on the client’s own hardware ensure data sovereignty, reduce latency to under 50ms for local inference, and eliminate per-token API costs at scale. The trade-off is higher initial infrastructure investment and the need for dedicated MLOps support for model updates. For a 201-500 employee firm, on-premise deployment typically requires 2-4 GPU servers and a dedicated MLOps engineer for the 6-month engagement. The model-agnostic architecture allows switching between cloud and on-premise models based on data sensitivity, with the same RAG pipeline and integration layer.

    3. RAG Over Confluence Turns Your Knowledge Base Into a Review Engine

    The RAG pipeline indexes contract templates, past executed agreements, and compliance checklists from Confluence or Notion into a vector database. When a new contract arrives, the system extracts key clauses (liability caps, SLA terms, termination conditions) and retrieves relevant precedents from the knowledge base. The LLM drafts a review summary highlighting deviations from standard terms, flagging clauses that exceed risk thresholds. Human reviewers approve or reject each flag before the contract proceeds to signature. The system logs every decision, creating an audit trail for compliance. Integration with the existing ERP ensures that approved contracts automatically update vendor master data and payment terms. The conversational agent handles initial intake, extracting metadata and routing contracts to appropriate reviewers based on risk classification, reducing ticket volume to legal by 40-60%.

    4. Human-in-the-Loop Approval Is Non-Negotiable for Money and Liability

    The most common failure is treating AI as a replacement for human judgment rather than an augmentation tool. Firms that remove human approval for contracts touching money, liability, or regulatory compliance face significant risk. The second pitfall is insufficient baseline measurement: without pre-implementation data on cycle time and error rate, you cannot prove ROI or identify where the AI is actually helping. The third is poor integration: if the AI assistant doesn’t plug into the existing CRM, ERP, and helpdesk via APIs, it creates a parallel workflow that increases rather than reduces manual work. The fourth is model selection mismatch: using cloud APIs for data that must stay on-premise, or using open-weight models when cloud quality is acceptable and cost-effective. Each pitfall has a measurable cost: unapproved AI decisions can trigger contract disputes, missing baselines make ROI unprovable, poor integration adds 20-30% overhead, and model mismatch increases costs by 40-60%.

    5. Dedicated AI Team Embeds in Your Org for the Full 6 Months

    A dedicated AI team typically includes a technical lead for architecture and model selection, a product designer for workflow mapping and human-in-the-loop UX, two full-stack developers for API integrations with ERP/CRM systems, and an MLOps engineer for on-premise model deployment and monitoring. For a 201-500 employee firm, this team operates as an embedded unit within the client’s organization for the 6-month engagement, with weekly steering meetings and bi-weekly demo cycles. The team size scales with complexity: a single contract type pilot requires 4-5 people, while multi-type rollout may expand to 6-8. Post-engagement, a subset (1-2 people) transitions to managed operation support. The dedicated team model ensures continuity: the same people who built the system understand its failure modes and can respond to edge cases within 4-8 hours, compared to 24-48 hours for external support contracts.

    6. Six-Month Timeline With Measurable Checkpoints at Each Phase

    The 6-month timeline breaks down as: Weeks 1-4 for process audit and baseline measurement of current cycle times and error rates. Weeks 5-10 for pilot development on one contract type (e.g., carrier agreements), including RAG pipeline setup and integration with Confluence/Notion. Weeks 11-16 for pilot validation, error rate measurement, and human-in-the-loop workflow refinement. Weeks 17-24 for rollout to additional contract types, team training, and managed operation handoff. Each phase includes measurable checkpoints: the pilot must demonstrate at least 30% cycle time reduction and 50% error rate improvement before rollout proceeds. The final deliverable is a fully operational AI contract review system integrated with existing ERP, CRM, and helpdesk, with a documented runbook for the internal team to manage day-to-day operations. The system is model-agnostic, allowing future migration to newer models without re-architecting the pipeline.

  • Three Months to Cut Back-Office Errors in a German Fintech

    1. Start with a Process Audit, Not a Model

    A German fintech with 30 employees processes 400 payment-related documents per week. The back-office team spends 12 hours a week manually extracting data from invoices and payment confirmations, with a 4% error rate that triggers reconciliation delays. Forfis starts with a process audit that maps every manual touchpoint, then selects document extraction as the pilot workflow. The fixed-scope pilot runs for six weeks, shipping with a measured baseline: cycle time drops from 18 minutes per document to 4 minutes, and the error rate falls to 0.8%. The pilot’s success criteria are explicit and tied to the audit’s findings, not vague “efficiency gains.”

    2. Run the Pilot on Document Extraction

    The pilot targets one workflow: extracting line items, amounts, and reference numbers from payment statements and invoices. Forfis uses an open-weight model on the client’s own hardware because PCI DSS requires cardholder data to stay within a controlled environment. The model runs on a single GPU server in the client’s Frankfurt data center. The extraction pipeline feeds directly into the existing ERP via API, so no new data store is introduced. A human reviews every extracted record before it posts to the ledger, satisfying the human-in-the-loop requirement for anything touching money.

    3. Layer a Lead-Qualification Assistant on the CRM

    With the back-office pilot validated, the second phase adds a customer-facing AI assistant for lead qualification. The assistant pulls from the CRM and a Confluence knowledge base to draft first-response emails for inbound leads. It classifies each lead by intent, budget range, and product fit, then flags high-value prospects for the sales team. A rep approves every outbound message before it sends. The assistant reduces initial qualification time from 25 minutes to under 5 per lead, and the sales team reports a 15% lift in response rate within the first month of rollout.

    4. Keep the Stack Model-Agnostic and On-Premise

    The architecture is deliberately model-agnostic. OpenAI and Anthropic APIs handle non-sensitive tasks like drafting marketing copy or summarizing meeting notes. Open-weight models on the client’s hardware handle anything touching payment data, health records, or contracts. This split lets the fintech use frontier models where quality matters most while keeping regulated data on-premise. The integration layer plugs into the existing CRM, ERP, and helpdesk through their native APIs, so no system is replaced. For a 30-person team, this means no new vendor lock-in and no migration project.

    5. Scale Across Departments in the Third Month

    After the pilot, the rollout extends to two adjacent departments: the finance team adopts the document extraction pipeline for vendor invoices, and the support team uses the same RAG assistant for ticket triage. The key is that each new workflow reuses the same architecture, the same on-premise model, and the same human-approval gate. Forfis ships a measured before/after baseline for every workflow: cycle time, error rate, and cost per transaction. By month three, the back-office error rate has dropped from 4% to 0.8% across all automated workflows, and the team has freed up roughly 20 hours per week for higher-value work.

    6. Ship a Measured Baseline, Not a Promise

    The three-month timeline works because the scope is fixed and the success criteria are measurable. The process audit takes two weeks, the pilot runs six weeks, and the rollout occupies the final four weeks. For a 30-person fintech in Germany, this means no open-ended engagement and no surprise invoices. The human-in-the-loop design means the team never has to trust the model blindly: anything touching money, contracts, or health data gets a human sign-off. The result is a back office that runs on 0.8% error rates, a sales team that responds to leads in under five minutes, and an architecture that keeps PCI DSS-compliant data on the client’s own hardware.

  • How a 15-Person UK Medtech Firm Cut Back-Office Errors 40% in Two Weeks

    1. The pilot scope is one workflow, not a platform

    A 15-person medtech company in Manchester was losing 11 hours per week to manual invoice data entry and document extraction. The finance lead typed supplier invoices into the ERP, cross-checked line items against purchase orders, and flagged discrepancies in a shared spreadsheet. Error rate: 6.2% on a sample of 200 invoices. Cycle time: 4.3 hours per batch.

    The fix was not a new hire. It was a fixed-scope pilot with a two-week deadline: automate the extraction and validation step for one supplier, integrate it into the existing ERP via API, and measure the before/after delta. The pilot used the Anthropic Claude API for document parsing because the invoice formats were inconsistent and required nuanced field mapping. A human approved every extracted record before it hit the ERP. The result: error rate dropped to 1.8%, cycle time fell to 1.1 hours per batch, and the finance lead spent the freed time on supplier negotiations instead of data entry.

    2. The integration lives in Slack, not a new dashboard

    The pilot ran inside the team’s existing Slack workspace. A bot posted extracted invoice fields into a dedicated channel, tagged the finance lead for approval, and logged the decision. No new UI, no new login, no training session. The integration used the Slack API and the ERP’s REST endpoint — both already in production.

    This matters because a 15-person team does not have the bandwidth to adopt a new tool. The workflow orchestration layer sat between the Claude API and the ERP: it handled retries, format validation, and the approval gate. When the finance lead approved a record in Slack, the orchestration layer pushed it to the ERP. When they rejected it, the bot asked for the correction and re-processed. Every interaction was logged for the GDPR audit trail. The team never left Slack. The AI never replaced the ERP. It filled the gap between the two.

    3. GDPR compliance is a design constraint, not an afterthought

    The pilot processed supplier invoices, which contain no patient data. But the company’s broader documentation — SOPs, regulatory checklists, clinical trial protocols — does. The architecture was designed from day one to be model-agnostic: the orchestration layer could route a request to the Anthropic Claude API for general document work, or to an open-weight model running on the company’s own server for anything touching special-category data under GDPR Article 9.

    The DPIA was completed before the pilot started. It documented: what data the AI processes, where it is stored, who can access it, and how a human can override any automated decision. The Data Processing Agreement with Anthropic was signed. The open-weight model (a 7B-parameter Llama variant) ran on a single GPU workstation in the office. No patient data left the building. The pilot’s scope was narrow enough that the compliance overhead was a one-day task, not a multi-week project.

    4. The baseline is measured, not assumed

    The pilot’s success metric was not “the AI works.” It was: error rate drops from 6.2% to under 3%, and cycle time drops from 4.3 hours to under 2 hours per batch. The baseline was measured in week one, before any automation was live. The team processed 50 invoices manually and logged every error and every minute. In week two, the AI processed the same 50 invoices, and the finance lead approved or corrected each one. The delta was the deliverable.

    This is what separates a pilot from a demo. A demo shows the AI extracting fields from a sample PDF. A pilot measures whether the extraction is accurate enough to trust in production, and whether the human approval step is fast enough to be worth the overhead. The 4.3-hour to 1.1-hour drop was not theoretical. It was logged in the ERP’s audit trail, timestamped, and attributable to the automation.

    5. The rollout is a sequence of fixed-scope engagements

    The pilot’s scope was one supplier, one document type, one integration point. The rollout plan was explicit: week three adds the second supplier, week four adds the third, week five adds the document extraction for purchase orders. Each expansion was a separate fixed-scope engagement with its own baseline and success metric.

    This is how a 15-person company scales operations without new hires. The finance lead’s role did not change — she still approved every record. But the time she spent typing dropped from 4.3 hours to 1.1 hours per batch. The freed capacity went to supplier management, which had been neglected for two years. The company did not hire a data entry clerk. It did not buy a new ERP. It added an AI layer to the workflow it already ran, measured the delta, and expanded only when the numbers justified it.

    6. The knowledge search is a byproduct, not the goal

    The pilot’s real value was not the 40% error reduction. It was the internal knowledge search capability that emerged from the same architecture. The orchestration layer that routed invoice data to the ERP was repurposed to route queries to the company’s document store. A support agent in Slack could now ask, “What is the recall procedure for device X?” and get a cited answer from the SOP, with the relevant section highlighted. The agent still reviewed the answer before sending it to a customer. The AI did not replace the agent. It cut the search time from 12 minutes to 90 seconds.

    The synthesis: a 15-person UK medtech company did not need a new hire, a new ERP, or a new helpdesk. It needed a two-week fixed-scope pilot that measured a real delta, ran inside the tools the team already used, and kept a human in the loop for every high-stakes action. The AI was a layer, not a replacement. The compliance was a constraint, not a blocker. The rollout was a sequence, not a big bang. That is the pattern that works when the team is small, the data is regulated, and the timeline is two weeks.