Background: A UK Medtech Distributor at 1,200 Headcount
This case study is a composite drawn from patterns Forfis has observed across multiple engagements. We do not name real clients. The company described here is a UK-based medtech distributor with roughly 1,200 employees, operating in the 501-2000 band. It handles procurement, supply-chain coordination, and customer-facing service for hospital and clinic clients across the UK and Ireland. The existing stack includes a mid-market ERP, a CRM for customer records, and Microsoft Teams as the primary internal messaging channel. The finance and operations teams were running on a mix of spreadsheets, email threads, and a legacy invoice portal that had not been updated since 2019. The company had no dedicated AI team and had not previously deployed any machine-learning system in production.
Challenge: 48-Hour Invoice Response, Zero New Hires, GDPR in the Loop
The trigger was a 40 percent increase in supplier invoice volume over eighteen months, driven by a new product line and expanded distribution contracts. The finance team of eleven was processing invoices manually: extracting line items, matching them against purchase orders, flagging discrepancies, and posting to the ERP. Average first-response time to a supplier query about a disputed invoice was 48 hours. The operations director had a hard constraint: no new headcount in the current fiscal year, and GDPR compliance was non-negotiable because invoice metadata occasionally contained patient-identifiable information from hospital procurement orders. The deadline was six months to show a measurable reduction in cycle time before the next board review. The team needed to cut first-response time without adding a single FTE and without sending regulated data to a third-party cloud API.
Approach: On-Premise Open-Weight Models, Predictive Scoring, and a Fixed-Scope Pilot
Forfis ran a two-week process audit across the finance and operations workflows. The audit identified invoice processing as the highest-impact target: high volume, repetitive extraction, and a clear before/after metric. The pilot scope was fixed: one invoice category (supplier purchase orders with line-item extraction), one integration point (the existing ERP API), and one notification channel (Microsoft Teams). The architecture used open-weight models on the client’s own hardware, so no regulated data left the building. A retrieval-augmented layer pulled context from the client’s own procurement documentation and CRM records to improve extraction accuracy. Predictive scoring assigned a confidence value to each extracted field; items above 95 percent auto-posted, items below routed to a human reviewer in Teams. The dedicated AI team of four engineers and one product designer worked on-site for the first four weeks, then shifted to remote with weekly syncs. The pilot ran for eight weeks with a measured baseline captured in week one.
Outcome: 48 Hours to Under 6, Error Rate Below 2 Percent
The pilot cleared its threshold. Average first-response time for supplier invoice queries dropped from 48 hours to under 6 hours. Extraction error rate on line items fell from 11 percent to under 2 percent. The finance team’s manual review volume dropped by roughly 60 percent, because the predictive scoring layer auto-approved the high-confidence items. The remaining 40 percent of invoices still required human eyes, but the reviewers now worked from a pre-drafted, context-enriched queue in Teams rather than a blank spreadsheet. The ERP integration held: no data left the client’s infrastructure, and the GDPR data-processing record was updated to reflect the on-premise model deployment. The operations director reported that the team absorbed the 40 percent invoice volume increase without a single new hire. The six-month timeline was met, and the board review proceeded on the strength of the measured baseline.
Lessons for Similar Teams
- Baseline first, always. The pilot did not start until the team had a measured before/after baseline on cycle time and error rate. Without that number, the board review would have been a conversation about impressions rather than data. Every similar team should capture the baseline in week one, not after the pilot ends.
- Model-agnostic architecture pays off. The client started with open-weight models on-premise for GDPR reasons. If a future use case requires a frontier API for a non-regulated workflow, the integration layer does not need to be rebuilt. Teams that hard-code a single vendor API into their architecture will face this problem.
- Predictive scoring is the human-in-the-loop mechanism. The confidence threshold is not a suggestion; it is the architectural gate. Items above 95 percent auto-approve, items below route to a human. This is what makes GDPR Article 22 compliance operational rather than theoretical.
- Integration through existing APIs, not replacement. The ERP, CRM, and Teams stack stayed intact. The AI layer sat on top. For a 1,200-person operation, a rip-and-replace project would have taken two years and a budget the company did not have.
- Dedicated team beats rotating contractors. The four engineers and one product designer stayed on the engagement from audit through rollout. Consistency in the team meant the client’s internal stakeholders had a single point of contact and a shared context that did not reset every sprint.
Leave a Reply