How a 15-Person UK Medtech Firm Cut Back-Office Errors 40% in Two Weeks

1. The pilot scope is one workflow, not a platform

A 15-person medtech company in Manchester was losing 11 hours per week to manual invoice data entry and document extraction. The finance lead typed supplier invoices into the ERP, cross-checked line items against purchase orders, and flagged discrepancies in a shared spreadsheet. Error rate: 6.2% on a sample of 200 invoices. Cycle time: 4.3 hours per batch.

The fix was not a new hire. It was a fixed-scope pilot with a two-week deadline: automate the extraction and validation step for one supplier, integrate it into the existing ERP via API, and measure the before/after delta. The pilot used the Anthropic Claude API for document parsing because the invoice formats were inconsistent and required nuanced field mapping. A human approved every extracted record before it hit the ERP. The result: error rate dropped to 1.8%, cycle time fell to 1.1 hours per batch, and the finance lead spent the freed time on supplier negotiations instead of data entry.

2. The integration lives in Slack, not a new dashboard

The pilot ran inside the team’s existing Slack workspace. A bot posted extracted invoice fields into a dedicated channel, tagged the finance lead for approval, and logged the decision. No new UI, no new login, no training session. The integration used the Slack API and the ERP’s REST endpoint — both already in production.

This matters because a 15-person team does not have the bandwidth to adopt a new tool. The workflow orchestration layer sat between the Claude API and the ERP: it handled retries, format validation, and the approval gate. When the finance lead approved a record in Slack, the orchestration layer pushed it to the ERP. When they rejected it, the bot asked for the correction and re-processed. Every interaction was logged for the GDPR audit trail. The team never left Slack. The AI never replaced the ERP. It filled the gap between the two.

3. GDPR compliance is a design constraint, not an afterthought

The pilot processed supplier invoices, which contain no patient data. But the company’s broader documentation — SOPs, regulatory checklists, clinical trial protocols — does. The architecture was designed from day one to be model-agnostic: the orchestration layer could route a request to the Anthropic Claude API for general document work, or to an open-weight model running on the company’s own server for anything touching special-category data under GDPR Article 9.

The DPIA was completed before the pilot started. It documented: what data the AI processes, where it is stored, who can access it, and how a human can override any automated decision. The Data Processing Agreement with Anthropic was signed. The open-weight model (a 7B-parameter Llama variant) ran on a single GPU workstation in the office. No patient data left the building. The pilot’s scope was narrow enough that the compliance overhead was a one-day task, not a multi-week project.

4. The baseline is measured, not assumed

The pilot’s success metric was not “the AI works.” It was: error rate drops from 6.2% to under 3%, and cycle time drops from 4.3 hours to under 2 hours per batch. The baseline was measured in week one, before any automation was live. The team processed 50 invoices manually and logged every error and every minute. In week two, the AI processed the same 50 invoices, and the finance lead approved or corrected each one. The delta was the deliverable.

This is what separates a pilot from a demo. A demo shows the AI extracting fields from a sample PDF. A pilot measures whether the extraction is accurate enough to trust in production, and whether the human approval step is fast enough to be worth the overhead. The 4.3-hour to 1.1-hour drop was not theoretical. It was logged in the ERP’s audit trail, timestamped, and attributable to the automation.

5. The rollout is a sequence of fixed-scope engagements

The pilot’s scope was one supplier, one document type, one integration point. The rollout plan was explicit: week three adds the second supplier, week four adds the third, week five adds the document extraction for purchase orders. Each expansion was a separate fixed-scope engagement with its own baseline and success metric.

This is how a 15-person company scales operations without new hires. The finance lead’s role did not change — she still approved every record. But the time she spent typing dropped from 4.3 hours to 1.1 hours per batch. The freed capacity went to supplier management, which had been neglected for two years. The company did not hire a data entry clerk. It did not buy a new ERP. It added an AI layer to the workflow it already ran, measured the delta, and expanded only when the numbers justified it.

6. The knowledge search is a byproduct, not the goal

The pilot’s real value was not the 40% error reduction. It was the internal knowledge search capability that emerged from the same architecture. The orchestration layer that routed invoice data to the ERP was repurposed to route queries to the company’s document store. A support agent in Slack could now ask, “What is the recall procedure for device X?” and get a cited answer from the SOP, with the relevant section highlighted. The agent still reviewed the answer before sending it to a customer. The AI did not replace the agent. It cut the search time from 12 minutes to 90 seconds.

The synthesis: a 15-person UK medtech company did not need a new hire, a new ERP, or a new helpdesk. It needed a two-week fixed-scope pilot that measured a real delta, ran inside the tools the team already used, and kept a human in the loop for every high-stakes action. The AI was a layer, not a replacement. The compliance was a constraint, not a blocker. The rollout was a sequence, not a big bang. That is the pattern that works when the team is small, the data is regulated, and the timeline is two weeks.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *