Background: A 300-Person Logistics Firm Stuck in Pilot Purgatory
This case study is a composite drawn from patterns Forfis has observed across multiple engagements in German logistics and supply-chain firms. No named customer is represented. The details below reflect a recurring profile: a mid-size operator in the 201-500 employee band, running on a legacy ERP, under pressure to scale without adding headcount, and sitting in the “running isolated pilots” stage of AI maturity. The company in this narrative is a fictional stand-in for that profile.
The firm, which we will call TransLog GmbH, operates a 300-person logistics and supply-chain business out of Frankfurt. It manages inbound freight for mid-market e-commerce brands and B2B distributors across DACH. Its stack is a mix of SAP Business One for finance and inventory, Notion as the internal knowledge base and project tracker, and a patchwork of spreadsheets and email for contract management. The finance and accounting team of 14 people handles invoice processing, carrier rate agreements, and vendor contracts manually. The CTO is a former operations lead who has approved two small AI experiments (a chatbot on the website, a spreadsheet macro for invoice categorization) but has not yet committed to a structured automation program. The company is in the running isolated pilots stage: it has tried AI, but the pilots never left the sandbox, and no one owns the rollout path.
Challenge: 4-Day Contract Review, Zero Headcount Budget
The trigger was a 40% volume increase in inbound carrier contracts over two quarters, driven by a new e-commerce client. The finance team was already at capacity: 14 people processing roughly 1,200 contracts and 4,500 invoices per month. The average first-response time for a new carrier rate agreement was 4 business days from receipt to validated entry in SAP. The error rate on liability-cap and indemnity fields was 3.2%, and each correction required a phone call to the carrier, adding 2-3 days of delay. The CFO had a hard deadline: the new client’s contract portfolio had to be fully onboarded by the end of Q3, and the board had frozen headcount for the year. The CTO’s ask was specific: cut first-response time on contract review without hiring, and keep the solution inside the existing stack. No new SaaS subscriptions, no data leaving the building for anything touching carrier financial terms. The EU AI Act was a secondary but non-negotiable constraint: the firm’s legal counsel had flagged that any AI system processing contracts with legal effect needed a documented human-oversight layer and a model-logging trail.
Approach: A Fixed-Scope Integration Sprint on n8n
Forfis ran a process audit in weeks 1-2, sampling 80 historical carrier rate agreements and timing the manual workflow. The audit confirmed the 4-day cycle and identified three bottleneck stages: PDF-to-text conversion (manual, 15 min per document), field extraction (manual, 25 min), and SAP entry (10 min). The pilot scope was fixed: one document type (carrier rate agreements), 14 extraction fields, one human-approval gate, and two integration endpoints (Notion for review, SAP for final write). The architecture used n8n as the orchestration layer: a webhook received the PDF from the shared drive, an OCR step converted it to text, an LLM call (OpenAI API for the initial extraction pass, with a fallback to an open-weight model on the client’s own hardware for fields containing financial terms) produced a structured JSON, and a confidence-score router sent low-confidence fields to a Notion review board. The human reviewer saw the original PDF page, the extracted value, and the model’s confidence score. Approved records were written back to SAP via its BAPI interface. The entire pipeline was built in weeks 3-6, tested in shadow mode against 200 historical documents in weeks 7-10, and went live in week 11 with a 2-week hypercare window.
Outcome: 94% Cycle-Time Reduction, 0.4% Error Rate
After the 2-week hypercare period, the measured results were as follows. Cycle time for a carrier rate agreement dropped from 4.1 business days to 6.2 hours, a 94% reduction. The 6-hour figure includes the human-approval step: the n8n pipeline processed the document in under 90 seconds, but the reviewer’s SLA was 4 hours, and the SAP write-back added 30 minutes. Error rate on the 14 extraction fields fell from 3.2% to 0.4%, with the remaining errors concentrated in two fields: the liability cap (0.8% error) and the force-majeure clause reference (0.3%). The finance team processed 1,350 contracts in the first full month post-go-live, up from 1,200, with no additional headcount. The EU AI Act compliance checklist was satisfied: every model call was logged with prompt version, model identifier, and confidence score in a read-only Notion database; the human-approval gate was documented in the firm’s AI governance policy; and the open-weight model for financial fields ran on the client’s own GPU server, so no regulated data left the building. The CFO’s Q3 deadline was met with 11 days to spare.
Lessons for Teams Running Isolated Pilots
- Fix the scope before you build. The pilot succeeded because the 14-field schema and the single document type were locked in week 1. Two scope changes were requested during the sprint (adding a force-majeure sub-field and a second document type); both were logged as change requests and deferred to a phase-2 sprint. Without that discipline, the 3-month timeline would have slipped to 5.
- Build the audit log from day one, not after go-live. The EU AI Act’s logging requirement (Article 12 for high-risk, Article 13 for transparency) is easier to satisfy when the n8n workflow writes every model call to a structured log from the first test run. Retrofitting logging after go-live forced a 3-day rework in one of Forfis’s other engagements.
- Set the human-approval SLA before the pipeline goes live. The 4-hour reviewer SLA was agreed with the finance team in week 2. Without it, the pipeline would have become a bottleneck: documents would have piled up in the Notion review board, and the cycle-time gain would have evaporated.
- Use the open-weight model for regulated fields, not as a cost-cutting default. The decision to run the financial-term extraction on the client’s own hardware was driven by the data-residency constraint, not by model quality. The OpenAI API handled the bulk extraction; the local model handled the sensitive fields. This split kept the architecture model-agnostic and the compliance story clean.
- Measure error rate per field, not as an aggregate. A 0.4% aggregate error rate sounds reassuring, but the 0.8% on the liability cap was the field that mattered. Reporting per-field errors in the weekly hypercare report kept the finance team’s trust and surfaced the one prompt that needed tuning.
Leave a Reply