Background: A 1,200-Person Medtech Firm in Graz
This case study is a composite drawn from patterns Forfis has observed across multiple engagements in the healthcare and medtech sector. It does not describe a single named client. The company, the metrics, and the timeline are representative of what we see in the field when a mid-sized European healthcare organization moves from isolated AI pilots to a compliance-safe, managed rollout of invoice-processing automation.
The company is a 1,200-person medtech firm based in Graz, Austria, manufacturing surgical instruments and diagnostic kits. It operates in 14 EU markets and reports under Austrian GAAP with quarterly IFRS reconciliation. The finance and accounting team is 42 people, of whom 11 handle accounts payable and monthly reporting. Their ERP is SAP S/4HANA, their document management system is a legacy on-prem archive, and their internal communication runs on Microsoft Teams. They had run two prior AI pilots — one for email triage, one for contract clause extraction — but neither had moved past the pilot stage. The finance director’s mandate was clear: automate the monthly invoice-to-reporting cycle without introducing a new compliance surface, and do it within a 4-week pilot window before the Q3 close.
The Challenge: 3,400 Invoices, 11 Staff, a 4-Week Window
The monthly reporting cycle ran from the 1st to the 10th of each month. During that window, 11 finance staff manually processed roughly 3,400 vendor invoices, extracted line items, matched them to purchase orders, flagged discrepancies, and posted entries to SAP. The cycle time from invoice receipt to ERP posting averaged 6.2 days, and the error rate — measured as manual corrections per 100 invoices — sat at 14.3. The finance director had a hard deadline: the Q3 close was in 11 weeks, and the board had asked for a visible efficiency gain by year-end. Headcount was not the constraint; the constraint was that the 11-person team could not absorb the 14% error rate without a second review pass, which doubled the cycle time. The prior two AI pilots had stalled because they were scoped as “AI projects” rather than as workflow replacements with a measured baseline. The finance director wanted a fixed-scope pilot with a before/after metric, not a proof of concept.
Approach: Audit, Then a Fixed-Scope Pilot on LangChain and LangGraph
Forfis ran a two-week AI Automation Audit before the pilot began. The audit mapped the invoice-to-reporting workflow end-to-end, sampled 200 invoices from the prior month, and established the baseline: 6.2-day cycle time, 14.3% error rate, 11 FTEs. The audit identified three automation points: (1) invoice ingestion and OCR extraction from the legacy archive, (2) line-item classification and PO matching, and (3) discrepancy flagging with a human approval step before ERP posting.
The pilot architecture used LangChain for the orchestration layer and LangGraph for the stateful workflow graph that tracked each invoice through ingestion, extraction, classification, review, and posting. The model layer was split: OpenAI’s GPT-4o handled the ambiguous line-item classification (where quality mattered), and an open-weight Llama 3 70B model ran on the client’s own GPU server for the regulated data extraction step, so no invoice data left the Graz data center. The integration points were SAP S/4HANA (via its OData API for posting approved entries), Microsoft Teams (for reviewer notifications and the approval workflow), and the legacy document archive (via a file-watcher ingestion pipeline). Every entry that touched money required a human approval in Teams before it hit SAP. The pilot ran for four weeks: Week 1 integration, Week 2 model tuning on the client’s actual invoice samples, Week 3 live operation with human-in-the-loop review, Week 4 measurement and the go/no-go report.
Outcome: 61% Cycle-Time Reduction, 3.8% Error Rate
By the end of Week 4, the pilot had processed 3,100 live invoices. The measured results against the audit baseline:
- Cycle time dropped from 6.2 days to 2.4 days (a 61% reduction). The bottleneck shifted from manual extraction to the human approval step, which the finance team chose to keep as a compliance control.
- Error rate fell from 14.3% to 3.8% (manual corrections per 100 invoices). The remaining errors were concentrated in two vendor categories with non-standard invoice formats, which the team flagged for a follow-up prompt-tuning pass.
- FTE allocation: the 11-person team redirected 4 FTEs from manual extraction to exception handling and vendor relationship management. The finance director did not reduce headcount; the freed capacity was absorbed into the Q3 close workload.
- Compliance surface: no new data left the building. The open-weight model ran on the client’s GPU server; the OpenAI API calls were limited to the classification step, which operated on anonymized line-item text, not on invoice metadata or vendor names.
The go/no-go report recommended a phased rollout to the remaining 12 EU markets over two quarters, with the same human-in-the-loop approval step retained for all money-touching entries.
Lessons for Teams Running Isolated Pilots
- Baseline before pilot, not after. The audit’s 200-invoice sample established the 6.2-day / 14.3% baseline before any model was tuned. Without that, the pilot’s results would have been unmeasurable. Teams that skip the baseline step cannot distinguish model improvement from natural variance.
- Split the model layer by data sensitivity, not by convenience. The open-weight model on the client’s hardware handled the regulated extraction step; the cloud API handled the classification step on anonymized text. This split is what made the rollout compliance-safe without requiring a full on-prem LLM deployment.
- Human-in-the-loop is a design constraint, not a fallback. The Teams approval step was in the LangGraph state machine from day one, not added after the pilot showed errors. Removing it post-hoc would have broken the workflow graph and required a re-architecture.
- Fixed-scope pilot, not open-ended POC. The 4-week window with a defined go/no-go report forced the team to ship a measurable result rather than iterate indefinitely. The finance director’s mandate — “show me a number by week 4” — was the single most important constraint in the engagement.
- Integration through existing APIs, not replacement. The system plugged into SAP, Teams, and the legacy archive through their native APIs. No system was replaced, which kept the rollout risk low and the change-management burden minimal.