Category: E-commerce and Retail

  • Five AI Workflow Patterns That Cut Manual Data Entry in E-Commerce

    1. Candidate Screening With Structured Extraction

    The highest-impact automation for a 501-2,000-person e-commerce firm is the candidate screening pipeline. Recruiters spend 40-60 minutes per resume manually extracting skills, experience, and education, then scoring against a rubric. A LangGraph-based workflow parses the resume PDF, extracts structured fields, scores against the job description, and flags edge cases for human review. The model drafts the screening summary; a recruiter approves or overrides. Cycle time drops from 45 minutes to 8 minutes per candidate, and the error rate in skill matching falls from 12% to 3% because the model is consistent and the human catches the remaining edge cases. This is the workflow that justifies the 6-month engagement because the volume is high, the manual steps are repetitive, and the before/after baseline is easy to measure.

    2. Invoice and PO Extraction Into the ERP

    E-commerce operations generate thousands of supplier invoices, purchase orders, and shipping documents per month. Manual data entry into the ERP is slow and error-prone. A document and data extraction pipeline uses an AI model to read the PDF or image, extract line items, totals, and vendor details, and write them to the ERP via API. The LangGraph orchestration handles the multi-step flow: parse, extract, validate against expected formats, flag low-confidence fields, and route to a human for approval if the confidence score is below threshold. For a mid-size retailer, this cuts invoice processing time by 60-70% and reduces data entry errors from 5% to under 1%. The human-in-the-loop step ensures that any invoice touching a financial record is approved by a person before it hits the general ledger.

    3. Orchestration Across Departments

    The first two workflows run in isolation. Workflow orchestration is what connects them into a coherent system. LangGraph models the state transitions: a candidate screening decision triggers a notification in Google Workspace, an invoice extraction flags a discrepancy that routes to the finance team’s inbox, and a document extraction error triggers a retry loop. The orchestration layer is model-agnostic, so the client can swap OpenAI for Anthropic or move to an open-weight model on their own hardware without re-architecting the workflow. For a 501-2,000-person firm, this means the AI team can add new workflows to the existing graph without rebuilding the integration layer. The dedicated team maintains the LangGraph state machine, monitors the approval queues, and tunes the model prompts based on the error rate data from the first 90 days.

    4. Google Workspace as the Human Interface

    The AI layer does not replace Google Workspace; it plugs into it. Screening summaries land in the recruiter’s Gmail inbox as structured emails. Invoice extraction results appear in a shared Google Drive folder with a summary sheet. Candidate rejection notifications go out via Google Calendar invites to schedule follow-ups. The integration uses the Google Workspace API, so the client’s existing authentication, permissions, and audit logs remain intact. For a mid-size e-commerce firm, this means the AI team does not need to build a new UI or force recruiters to adopt a new tool. The workflow is invisible: the recruiter opens their inbox, sees the AI-drafted screening summary, approves or edits it, and moves on. The before/after baseline tracks the time from resume receipt to recruiter decision, and the Google Workspace integration is what makes that measurement possible without adding a new system.

    5. Scaling the Pattern Across the Organization

    The pilot runs on one workflow, one department, one team. Scaling across departments means replicating the pattern: audit the next workflow, set the baseline, ship the pilot, measure the before/after, and roll out. For a 501-2,000-person e-commerce firm, the sequence is typically candidate screening (HR), then invoice processing (finance), then customer ticket triage (support), then document extraction for legal and compliance (contracts, NDAs, vendor agreements). Each workflow gets its own LangGraph state machine, its own human-in-the-loop approval queue, and its own baseline metrics. The dedicated AI team manages the rollout, tunes the models based on the error rate data, and ensures that the integration layer (Google Workspace, ERP, ATS) stays consistent across departments. The 6-month timeline covers the first two workflows; the remaining two follow in months 7-12.

  • UK E-Commerce Retailer Cuts Monthly Reporting from 14 Days to 36 Hours

    Background: A 1,200-Person UK E-Commerce Retailer

    This case study is a composite drawn from patterns Forfis has observed across multiple e-commerce and retail engagements in the UK. No named customer appears. The company described here is a mid-market online retailer with roughly 1,200 employees, operating across three fulfilment centres in the Midlands and the North of England. It sells through its own website and two major marketplaces, processes around 40,000 supplier invoices per month, and runs a monthly operations report that feeds into board-level KPIs. The existing stack includes a mid-tier ERP, a legacy document management system, and Microsoft Teams as the primary internal communication channel. The finance and operations teams are separate, and the monthly report is a hand-built spreadsheet assembled from exports in three different formats.

    The Challenge: 14 Days of Manual Reporting

    The monthly operations report took the finance team 14 working days to assemble. The process started with exporting supplier invoices from the document management system, manually keying line items into a spreadsheet, reconciling them against the ERP purchase orders, and then formatting the output for the board pack. Two analysts spent roughly 60 hours per cycle on this task, and the error rate on manual data entry sat around 4 to 6 percent, meaning roughly 1,600 to 2,400 line items per month required correction before the report could be signed off. The operations team, meanwhile, had no real-time visibility into supplier performance because the data was locked in the spreadsheet until the report was published. The pressure was not regulatory; it was operational. The CFO had flagged the reporting lag in a board review, and the head of operations wanted supplier scorecards available within 48 hours of month-end close, not 14 days later.

    Approach: Audit, Pilot, and n8n Orchestration

    Forfis began with a two-week AI automation audit. The audit mapped the invoice-to-reporting flow end to end, identified 11 distinct manual touchpoints, and scored each on volume, error rate, and cycle time. The top candidate was the invoice extraction and reconciliation step, which accounted for 70 percent of the analyst hours. The pilot scope was fixed at eight weeks: build a document and data extraction pipeline that ingests supplier invoices from the document management system, extracts line items, PO references, and tax codes, and pushes structured data into the ERP via its REST API. On top of that, a retrieval-augmented knowledge assistant was built over the company’s operations documentation, historical reports, and CRM records, accessible through Microsoft Teams. The orchestration layer was n8n, self-hosted on the client’s own infrastructure, so no data transited a third-party SaaS boundary. The model layer used OpenAI’s API for extraction quality and an open-weight model for the RAG assistant, running on the client’s GPU server, because the operations documentation contained supplier contract terms that procurement wanted to keep on-premises.

    Outcome: 36 Hours, Not 14 Days

    The pilot shipped in seven and a half weeks, one day ahead of the eight-week deadline. The extraction pipeline processed 40,000 invoices per month with a field-level accuracy of 96.2 percent on the test set, up from the 94 to 96 percent baseline of manual entry. The monthly report cycle dropped from 14 working days to 36 hours: the pipeline ran overnight, the RAG assistant generated a draft narrative summary by 09:00 the next morning, and a finance analyst reviewed and approved the output by 12:00. The error rate on the final report fell to under 1 percent. The operations team gained access to supplier scorecards within 48 hours of month-end close, a 12-day improvement. The two analysts who previously spent 60 hours per cycle on this task were redeployed to supplier negotiation support. The n8n workflow was handed over with documentation, and the client’s own operations team could adjust routing rules without a developer. The RAG assistant was scoped to the indexed corpus only; it did not have internet access, and access was controlled at the Teams channel level.

    Lessons for Similar Teams

    • Fix the pilot scope before writing code. The eight-week timeline held because the audit deliverable defined exactly which invoices, which fields, and which ERP endpoints were in scope. Any new request during the pilot was treated as a change order with its own timeline, not a silent addition. Teams that skip this step routinely blow past their deadline by two to three weeks.
    • Self-host the orchestration layer when procurement asks where data lives. n8n on the client’s own infrastructure answered that question in one sentence. A managed SaaS orchestrator would have required a data processing agreement and a security review that added three to four weeks to the timeline.
    • Partition the RAG index by department. The operations assistant could not query finance data, and vice versa. This was enforced at the vector store level, not just at the Teams channel level. Without partitioning, a user in logistics could have pulled supplier contract terms from the finance index.
    • Log every human approval with a timestamp and user ID. Even though no regulation mandated it, the audit trail became the first thing the CFO asked for in the post-pilot review. The log showed exactly who approved the report, when, and what the model had drafted before approval.
    • Model-agnostic from day one. The client swapped the RAG model from OpenAI to the open-weight model in week three when procurement raised a data-residency concern. The n8n workflow did not change; only the model endpoint did. That swap cost two hours of configuration, not a re-architecture.
  • UK E-commerce Voice Agent for Ticket Triage: 3-Month GDPR-Compliant Pilot

    Process Audit and Baseline Measurement

    A 500-person e-commerce company in the UK handles 12,000 support tickets monthly, with 40% involving order status checks or returns. The support team spends 6 hours per day on manual data entry and routing, with an average cycle time of 4.2 hours from ticket creation to first response. The goal is to reduce manual back-office work by 30% and cut cycle time to under 2 hours, while maintaining GDPR compliance and supporting English plus two additional languages. The engagement starts with a two-week process audit that analyzes call recordings, ticket logs, and CRM data to identify the top five query types and the current error rate. The audit produces a baseline document with cycle time, error rate, and customer satisfaction scores for each query type, which becomes the success criteria for the pilot. The team selects one product category and one language for the isolated pilot, ensuring the scope is fixed and measurable. The pilot runs for four weeks, with a human-in-the-loop approval for any action that touches money or account changes. The architecture uses the Anthropic Claude API for response generation, with a custom REST API and webhooks connecting the voice agent to the existing CRM and helpdesk. No data is stored in the AI layer; all records remain in the client’s systems. The pilot ships with a measured before/after baseline, and the team reviews the results in a structured debrief before deciding on rollout.

    Voice Agent Architecture and Model Selection

    The voice agent uses a three-stage pipeline: speech-to-text, language model inference, and text-to-speech. The speech-to-text engine captures the caller’s voice and converts it to text with a 92% accuracy rate in English. The Anthropic Claude API generates the response using a prompt template that includes the caller’s intent, order details, and the company’s returns policy. The prompt is tuned for each language, with a glossary of product terms and a confidence threshold that routes low-confidence calls to human agents. The text-to-speech engine converts the response to natural-sounding audio with a 180 ms latency, which is within the acceptable range for conversational AI. The system supports English, German, and French, with a fallback to English if the confidence score drops below 85%. The voice agent does not make decisions with legal or similar significant effects; it provides information and captures data, with a human agent handling any action that touches money or account changes. The architecture is model-agnostic, so the team can switch to an open-weight model on client hardware if the data sensitivity requires it. The integration layer uses custom REST APIs and webhooks to connect the voice agent to the CRM and helpdesk, with no vendor lock-in on the AI model or integration layer.

    Integration Sprint and API Design

    The integration sprint delivers a working voice agent connected to the client’s CRM and helpdesk via REST APIs and webhooks. The deliverable includes the model configuration, prompt templates, API endpoints, and a runbook for the support team. The client retains full ownership of the code and configuration, with no vendor lock-in on the AI model or integration layer. The integration layer is designed to be modular, so the team can add new languages or product categories without re-architecting the system. The API endpoints are documented with OpenAPI 3.0, and the webhooks are signed with HMAC-SHA256 to ensure data integrity. The system logs all interactions with a timestamp, caller ID, and intent classification, which the support team can query via the CRM’s reporting dashboard. The runbook includes troubleshooting steps for common issues, such as high latency or low confidence scores, and a contact list for the integration team. The client’s IT team is trained on the system during the final week of the sprint, with a handover document that covers the architecture, configuration, and maintenance procedures. The integration sprint is fixed-scope, with a defined deliverable and a 48-hour rollback window if the pilot fails to meet the success criteria.

    Isolated Pilot and Rollback Strategy

    The pilot runs in isolation on a single product category and one language, with a measured baseline of cycle time and error rate before go-live. The system does not touch production data or affect other support channels. If the pilot fails to meet the predefined success criteria, the team rolls back to the manual process within 48 hours, with no data loss or system disruption. The success criteria include a 30% reduction in manual data entry, a cycle time under 2 hours, and an error rate below 5%. The team reviews the results in a structured debrief, with a focus on the error types and the customer satisfaction scores. The debrief produces a report that includes the before/after metrics, the error analysis, and a recommendation for rollout. The rollout plan includes a phased approach, with the voice agent expanding to additional languages and product lines over the next eight weeks. The team monitors the error rate and customer satisfaction scores during the rollout, with a 24-hour review window where a support lead audits a sample of AI-handled calls for accuracy. The rollout is considered successful if the error rate remains below 5% and the customer satisfaction score does not drop by more than 2 points.

    GDPR Compliance and Data Handling

    The system complies with GDPR Article 5 (data minimization) and Article 22 (automated decision-making). Voice recordings and transcripts are encrypted in transit and at rest, with a lawful basis for processing. The data is stored in the client’s CRM and helpdesk, not in the AI layer, which reduces the data footprint and simplifies the compliance review. The team documents the logic of the AI system in a Data Protection Impact Assessment, which is required if the system makes decisions with legal or similar significant effects. The voice agent does not make such decisions; it provides information and captures data, with a human agent handling any action that touches money or account changes. The system offers a human review option for any caller who requests it, and the team maintains a log of all human reviews. The data retention policy is aligned with the client’s existing GDPR compliance program, with a maximum retention period of 12 months for voice recordings and 24 months for transcripts. The team conducts a quarterly review of the data processing activities, with a focus on the error rate and the customer satisfaction scores. The compliance review is documented in a report that is shared with the client’s data protection officer.

    Risk Mitigation and Error Handling

    The main risk is the voice agent providing incorrect information about order status or returns policy. Mitigation includes a human-in-the-loop approval for any action that touches money or account changes, a confidence threshold that routes low-confidence calls to humans, and a 24-hour review window where a support lead audits a sample of AI-handled calls for accuracy. The team monitors the error rate and the customer satisfaction scores during the pilot and rollout, with a focus on the error types and the root causes. The error analysis is documented in a report that is shared with the support team, with a focus on the corrective actions and the preventive measures. The team conducts a monthly review of the system’s performance, with a focus on the cycle time, the error rate, and the customer satisfaction scores. The review produces a report that includes the metrics, the error analysis, and a recommendation for improvement. The team maintains a knowledge base of common issues and their solutions, which is updated monthly based on the error analysis. The knowledge base is used to train the support team and to improve the prompt templates for the voice agent.

  • OpenAI API vs. Open-Weight Models for Invoice Extraction in Austrian E-Commerce

    What Is Being Compared

    The two options under comparison are the OpenAI API (specifically the gpt-4o-mini and gpt-4o models, accessed via HTTPS) and an open-weight model deployed on the client’s own hardware (Llama 3 70B or Mistral 8x7B, running on a single A100 80GB or a pair of L40S GPUs). Both options sit inside the same surrounding architecture: a document ingestion layer that pulls PDFs and scanned images from the ERP or email, an extraction pipeline that calls the model, a human-in-the-loop approval step, and an integration layer that posts the validated data back into SAP or Microsoft Dynamics. The model-agnostic design means the client can switch between the two options without rewriting the ingestion, approval, or integration code. The comparison below isolates the model layer and judges it against the eight criteria that matter for a 201-500 employee e-commerce operation in Austria running a 6-month engagement.

    Criteria for Judgment

    The following eight criteria frame the comparison. Each is chosen because it directly affects the 6-month timeline, the PCI DSS compliance posture, or the operational cost of scaling invoice processing across departments in an Austrian e-commerce firm.

    • Inference latency — measured from document submission to structured output, excluding human review time.
    • Per-document cost — API token fees or amortized GPU hardware cost per 1,000 invoices.
    • Data residency — whether document content leaves the client’s network boundary.
    • PCI DSS alignment — ease of meeting Requirement 3.4 (PAN rendering unreadable) and Requirement 10 (audit logging).
    • Integration effort — weeks required to connect the model layer to SAP or Dynamics via native API.
    • Vendor lock-in — cost and effort to switch to a different model provider after the pilot.
    • Compliance audit trail — whether the model provider retains logs that satisfy Austrian data-protection expectations under GDPR Article 30.
    • Scalability ceiling — maximum documents per day before the architecture requires a redesign.

    Side-by-Side Comparison

    Criterion OpenAI API (gpt-4o-mini) Open-Weight Model (Llama 3 70B on A100)
    Inference latency 1.2-2.8 s per invoice (p95) 0.8-1.5 s per invoice (p95)
    Per-document cost (1,000 invoices) USD 0.40-0.80 EUR 0.05-0.15 (amortized GPU)
    Data residency Documents transit OpenAI’s US/EU data centers All data stays on client’s on-prem hardware
    PCI DSS alignment Requires PAN tokenization before API call; OpenAI does not store data by default (zero-data-retention agreement available) No external transmission; PCI DSS scope limited to client’s own network
    Integration effort 2-3 weeks (HTTPS call, JSON response) 4-6 weeks (GPU provisioning, model serving stack, API gateway)
    Vendor lock-in Low; prompt and schema are portable Low; model weights are open, but serving stack is tied to specific hardware
    Compliance audit trail OpenAI provides request logs under ZDR agreement; client must maintain own logs for GDPR Art. 30 Full local logging; no third-party retention
    Scalability ceiling ~50,000 documents/day on a single API key ~8,000-12,000 documents/day on a single A100; linear scaling with additional GPUs

    When the OpenAI API Wins

    The OpenAI API wins when the 6-month timeline is the binding constraint. The 2-3 week integration effort versus 4-6 weeks for the open-weight path means the API option delivers a working pilot 3-4 weeks earlier, which is significant when the engagement must close within 26 weeks. For an Austrian e-commerce firm processing 500-2,000 supplier invoices daily, the API cost of USD 200-1,600 per month is a small fraction of the labor cost it replaces. The PCI DSS risk is manageable: invoices rarely contain PAN, and the zero-data-retention agreement with OpenAI eliminates the third-party retention concern. The API option also scales to 50,000 documents per day without hardware changes, which covers the scaling-across-departments scenario where the operations team later adds purchase orders, delivery notes, and credit memos to the same pipeline.

    The open-weight model wins when the compliance review explicitly forbids external data transmission. If the firm’s PCI DSS assessor or data-protection officer determines that even tokenized document content cannot leave the building, the on-prem path is the only option. The 4-6 week integration effort is absorbed by the 6-month timeline if the process audit starts in week 1 and the pilot begins in week 7. The per-document cost is lower at scale, but the upfront GPU hardware cost of EUR 10,000-15,000 (or EUR 2,000-3,000 per month rented) is a real budget line that the API option avoids.

    Recommendation for the 6-Month Engagement

    For a 201-500 employee e-commerce and retail firm in Austria running a 6-month engagement focused on invoice processing with SAP or Microsoft Dynamics integration, the OpenAI API is the recommended option. The rationale is threefold. First, the 2-3 week integration effort preserves 3-4 weeks of buffer within the 26-week timeline, which is critical because the process audit and baseline measurement phase often overruns by 1-2 weeks. Second, the PCI DSS risk is low for invoice processing: supplier invoices do not contain PAN, and the zero-data-retention agreement addresses the data-residency concern. Third, the scalability ceiling of 50,000 documents per day covers the scaling-across-departments scenario without a hardware redesign. The open-weight model remains the correct fallback if the compliance review in weeks 4-6 explicitly forbids external transmission, but that outcome is uncommon for invoice processing in e-commerce. The model-agnostic architecture ensures the client can switch to the open-weight path in 2-3 weeks if the compliance decision changes, without losing the pilot’s measured baseline.

  • Swiss E-commerce Cuts Invoice Cycle Time 92% in a Two-Week ISO 27001-Safe Pilot

    Background: A Swiss Retail Group Under Audit Pressure

    This case study is a composite drawn from patterns observed across multiple engagements. No named customer is represented. The details are plausible and reflect the range of outcomes seen in the field, but they do not describe a single real company.

    The client is a Swiss e-commerce and retail group with roughly 2,400 employees, operating in German, French, and Italian markets. The finance and accounting team handles 18,000 to 22,000 supplier invoices per month across three ERP instances. The stack is a mix of SAP S/4HANA for the core ledger, a legacy document management system for invoice images, and Confluence for internal runbooks and audit documentation. The company holds ISO 27001 certification and is in the middle of a renewal audit. The finance director’s mandate was clear: reduce the average cycle time from invoice receipt to ERP posting without introducing a compliance gap.

    Challenge: 20,000 Invoices a Month and a 90-Day Audit Clock

    The finance team was processing invoices manually: a clerk downloaded the PDF, typed the vendor name, amount, tax code, and cost center into the ERP, and flagged discrepancies for review. The average cycle time was 4 to 6 hours per invoice, with a 3 to 5 percent error rate on a sample of 500 invoices. The error rate was not just a cost issue; it was a compliance issue. ISO 27001 requires documented controls over financial data, and a 4 percent error rate on 20,000 invoices per month meant roughly 800 mis-posted entries that had to be caught in a secondary review. The secondary review was itself a manual process, adding another 2 to 3 hours per flagged invoice. The finance director had a deadline: the ISO 27001 renewal audit was 90 days out, and the auditor had already flagged the manual process as a control weakness.

    Approach: A Two-Week Pilot on the Top Five Vendors

    The engagement started with a three-day process audit. The team mapped the invoice lifecycle from receipt to posting, identified the 12 vendor categories that accounted for 78 percent of volume, and pulled a historical sample of 1,200 invoices for calibration. The pilot scope was fixed: one ERP instance, one vendor category (the top 5 suppliers by volume), and a two-week window. The architecture used the OpenAI API for extraction, with a human-in-the-loop approval queue. The model extracted vendor name, invoice number, amount, tax code, and cost center. A reviewer saw the proposed entry alongside the original PDF and could approve, correct, or reject. The approval log was written to Confluence and to the ERP audit trail. The pipeline connected to the ERP via its REST API and to the document store via SFTP. No new infrastructure was required. The client’s existing IT team handled the API credentials and network access.

    Outcome: 92 Percent Cycle-Time Reduction in 12 Days

    The pilot ran for 12 business days. The model processed 1,840 invoices from the top five vendors. The average cycle time dropped from 4.2 hours to 22 minutes, a 92 percent reduction. The error rate on the pilot sample was 0.8 percent, down from the 3.4 percent baseline. Of the 1,840 invoices, 1,612 were approved with zero edits. The remaining 228 required human correction, mostly on tax codes for cross-border invoices. The approval queue averaged 14 minutes per invoice for the corrected entries. The ISO 27001 audit trail showed 100 percent of inferences logged with timestamp, user ID, and confidence score. The finance director presented the pilot results to the audit committee. The auditor accepted the AI-assisted workflow as a control improvement, conditional on the managed operations SLA being in place before the renewal audit.

    Lessons for Teams Running Similar Pilots

    • The historical sample matters more than the model. The 1,200-invoice calibration sample was the single biggest factor in the 0.8 percent error rate. A team that skips this step and goes live with a generic prompt will see error rates of 8 to 12 percent and lose the human trust needed for the approval workflow.
    • Fix the scope before you start. The two-week window only worked because the pilot was limited to one ERP instance and five vendors. A team that tries to cover all 12 vendor categories in two weeks will spend the time on integration edge cases and miss the baseline measurement.
    • The approval queue is the product, not the model. The model’s extraction quality was good, but the reviewer interface was what made the workflow usable. A team that ships a model without a clean approval UI will see reviewers bypass the system and go back to manual entry.
    • ISO 27001 is a design constraint, not a post-hoc checkbox. The audit trail, the data processing agreement, and the access controls were built into the architecture from day one. Retrofitting them after go-live is 3 to 4 times more expensive and often fails the audit.
    • Managed operations is where the value compounds. The pilot proved the concept. The managed operations SLA, with monthly reports on confidence distribution and error rate, is what keeps the error rate at 0.8 percent instead of drifting to 3 percent as vendor formats change.
  • Swiss E-Commerce Firm Cuts Invoice Processing Time 71% with On-Premise AI

    Background: A 300-Person Swiss E-Commerce Firm at Capacity

    This case study is a composite based on patterns observed across Forfis engagements. We do not name real clients. The company described here is a mid-size e-commerce and retail operator based in Zurich, with roughly 300 employees across operations, customer service, and finance. The stack is a mix of a legacy ERP (SAP Business One), a modern CRM (HubSpot), and Slack as the primary internal communication channel. The company had already automated one process — a basic rules-based invoice matching workflow — and was looking to extend AI automation to the next layer of back-office work without adding headcount. The constraint was clear: the finance team was at capacity, and the CTO had a hard deadline to reduce manual data entry before the next fiscal year close.

    Challenge: 12 Hours a Week Lost to Manual Data Entry

    The finance team was spending an estimated 12 hours per week on manual document extraction: pulling supplier invoice fields (vendor name, amount, tax code, line items) from PDFs and entering them into the ERP. The error rate on manual entry was around 8%, and each correction cycle added 45 minutes of rework. The operational pressure was threefold: the fiscal year close was eight weeks away, the team had no budget for additional hires, and the company was in the middle of a PCI DSS re-certification audit, which meant any new system touching payment-related data had to pass a formal risk assessment under Requirement 12.8. The CTO needed a solution that would free senior staff from routine work without introducing a new compliance liability.

    Approach: On-Premise Llama 3 with a Slack Approval Loop

    Forfis ran a two-week process audit that mapped every manual touchpoint in the invoice processing workflow. The audit identified that 70% of the extraction work involved supplier invoices in a consistent PDF format, making them a strong candidate for a fixed-scope pilot. The pilot used an open-weight model (Llama 3 70B) fine-tuned on 500 historical invoice examples, running on the client’s own A100 GPU node inside their VPC. The integration layer connected to Slack: the AI posted extracted fields to a dedicated channel, a human approved or flagged each entry, and approved fields were pushed to the ERP via its REST API. The entire pilot ran in eight weeks, with a measured baseline captured in week one and a shadow run in weeks seven and eight.

    Outcome: 71% Faster Cycle Time, 2.4% Error Rate

    The pilot reduced the average cycle time per invoice from 14 minutes to 4 minutes, a 71% improvement. The field-level error rate dropped from 8% to 2.4%, below the 3% threshold agreed in the pilot scope. The human approval step required intervention on roughly 15% of documents in the first two weeks, tapering to 6% by the end of the shadow run. The finance team reported that the senior staff who had been doing manual entry were now spending that time on supplier negotiations and exception handling. The PCI DSS risk assessment was completed in week six, and the audit trail (every extraction event logged with a document hash) satisfied Requirement 10.2.2 without additional controls.

    Lessons for Teams Scaling AI Without New Hires

    • Baseline before you build. Capturing a 200-document baseline in week one is non-negotiable. Without it, you cannot prove the pilot worked, and the go/no-go decision becomes a gut call. Forfis treats the baseline as a contract: the same sample size, the same measurement method, before and after.
    • Pick the highest-volume, lowest-complexity workflow first. The pilot should target the workflow where the ratio of document volume to format variability is highest. A consistent PDF format with 70% of the volume is a better pilot candidate than a mixed-format pipeline with 30% of the volume.
    • The approval loop is the product, not the model. The Slack channel where a human clicks approve is where the real value lives. The model is a swappable component; the approval workflow is what the team actually uses every day.
    • PCI DSS compliance is a design constraint, not an afterthought. The on-premise architecture and the audit trail were built in from day one, not bolted on after the pilot. Requirement 12.8 risk assessment and Requirement 10.2.2 logging were part of the pilot scope, not a separate workstream.