Background: A 2,400-Person Frankfurt Firm Stuck in Pilot Purgatory
This case study is a composite drawn from patterns observed across multiple engagements. No named customer appears here; the details are aggregated and anonymized to protect client confidentiality. The firm in question is a 2,400-person professional services company based in Frankfurt, operating across legal, tax, and consulting practices. It runs a mid-sized ERP, a Confluence instance for internal documentation, and a shared inbox for incoming invoices. The finance team of 38 people handled roughly 12,000 invoices per month, with a manual cycle time of 4.2 days from receipt to posting. The firm had run two prior AI pilots, both isolated and both abandoned after the pilot phase ended. It was in the “running isolated pilots” stage of AI maturity: the technology was proven in small tests, but no workflow had crossed the threshold into production.
Challenge: 12,000 Invoices a Month, 38 People, and a Year-End Close
The finance director’s mandate was specific: cut the first-response time on invoice processing without adding headcount. The operational pressure was a combination of a year-end close deadline, a 12 percent increase in invoice volume from two new client engagements, and a two-person vacancy in the accounts payable team. The firm had no compliance constraints beyond standard German tax law, but the finance team was risk-averse: any system that touched a bank transfer or a contract clause required a human approval step. The prior pilots had failed because they were open-ended, lacked a measured baseline, and did not integrate with the existing ERP. The team needed a fixed-scope engagement with a clear success metric and a handover plan that did not lock them into a vendor subscription.
Approach: LangGraph Workflow, Model-Agnostic Architecture, and a Human Approval Queue
Forfis ran an eight-week fixed-scope pilot on the invoice processing workflow. The architecture was model-agnostic: OpenAI’s GPT-4o handled the extraction and classification steps, while an open-weight Llama 3 model on the client’s own hardware processed the sensitive fields that could not leave the building. The orchestration layer was LangGraph, which managed the state machine for the extraction, validation, and approval steps. The system ingested PDFs and scanned images from the ERP, extracted line items, tax codes, vendor names, and payment terms, then cross-checked them against the purchase order. If the confidence score was above the threshold, it posted the entry automatically; if not, it routed the invoice to a human reviewer in a queue. The integration used the ERP and Confluence APIs, not a new platform. The runbook and monitoring dashboard were part of the deliverable.
Outcome: 42 Percent Faster Cycle Time, 55 Percent Fewer Errors
The pilot met both success criteria by week six. The average cycle time dropped from 4.2 days to 2.4 days, a 42 percent reduction. The error rate on manual entries fell from 3.1 percent to 1.4 percent, a 55 percent cut. The approval queue depth stayed under 15 invoices at any given time, which the finance team found manageable. The system handled 94 percent of invoices without human intervention; the remaining 6 percent were routed to the queue, where the average review time was 11 minutes per invoice. The finance team reported that the Confluence updates for vendor payment history were accurate and useful, and the monitoring dashboard gave them visibility into the confidence scores and error trends. The year-end close was completed on schedule, with the finance team reporting that the system absorbed the 12 percent volume increase without additional headcount.
Lessons for Teams Running Isolated Pilots
- Measure the baseline before you build. The team tracked cycle time and error rate for two weeks before the pilot started. Without that baseline, the 42 percent improvement would have been anecdotal rather than defensible. The success criteria were agreed in week one and not reopened mid-flight.
- Model-agnostic from day one. The LangGraph workflow was designed so that swapping OpenAI for an open-weight model was a configuration change, not a rewrite. This mattered when the client’s security team flagged that certain vendor fields could not leave the building.
- The approval queue is the product, not the model. The finance team’s trust in the system came from the queue, not from the extraction accuracy. The queue was integrated with their existing task management tool, so they did not have to learn a new interface.
- Fixed scope is a feature, not a limitation. The eight-week timeline and the single workflow kept the team focused. The client did not ask for feature creep because the success criteria were clear and the handover plan was part of the deliverable.
- The runbook is the handover. The monitoring dashboard, the threshold tuning guide, and the escalation path were documented in the runbook. The client’s finance team could operate the system without Forfis on the phone.
Leave a Reply