Tag: Automate Monthly Reporting

  • UK E-Commerce Retailer Cuts Monthly Reporting from 14 Days to 36 Hours

    Background: A 1,200-Person UK E-Commerce Retailer

    This case study is a composite drawn from patterns Forfis has observed across multiple e-commerce and retail engagements in the UK. No named customer appears. The company described here is a mid-market online retailer with roughly 1,200 employees, operating across three fulfilment centres in the Midlands and the North of England. It sells through its own website and two major marketplaces, processes around 40,000 supplier invoices per month, and runs a monthly operations report that feeds into board-level KPIs. The existing stack includes a mid-tier ERP, a legacy document management system, and Microsoft Teams as the primary internal communication channel. The finance and operations teams are separate, and the monthly report is a hand-built spreadsheet assembled from exports in three different formats.

    The Challenge: 14 Days of Manual Reporting

    The monthly operations report took the finance team 14 working days to assemble. The process started with exporting supplier invoices from the document management system, manually keying line items into a spreadsheet, reconciling them against the ERP purchase orders, and then formatting the output for the board pack. Two analysts spent roughly 60 hours per cycle on this task, and the error rate on manual data entry sat around 4 to 6 percent, meaning roughly 1,600 to 2,400 line items per month required correction before the report could be signed off. The operations team, meanwhile, had no real-time visibility into supplier performance because the data was locked in the spreadsheet until the report was published. The pressure was not regulatory; it was operational. The CFO had flagged the reporting lag in a board review, and the head of operations wanted supplier scorecards available within 48 hours of month-end close, not 14 days later.

    Approach: Audit, Pilot, and n8n Orchestration

    Forfis began with a two-week AI automation audit. The audit mapped the invoice-to-reporting flow end to end, identified 11 distinct manual touchpoints, and scored each on volume, error rate, and cycle time. The top candidate was the invoice extraction and reconciliation step, which accounted for 70 percent of the analyst hours. The pilot scope was fixed at eight weeks: build a document and data extraction pipeline that ingests supplier invoices from the document management system, extracts line items, PO references, and tax codes, and pushes structured data into the ERP via its REST API. On top of that, a retrieval-augmented knowledge assistant was built over the company’s operations documentation, historical reports, and CRM records, accessible through Microsoft Teams. The orchestration layer was n8n, self-hosted on the client’s own infrastructure, so no data transited a third-party SaaS boundary. The model layer used OpenAI’s API for extraction quality and an open-weight model for the RAG assistant, running on the client’s GPU server, because the operations documentation contained supplier contract terms that procurement wanted to keep on-premises.

    Outcome: 36 Hours, Not 14 Days

    The pilot shipped in seven and a half weeks, one day ahead of the eight-week deadline. The extraction pipeline processed 40,000 invoices per month with a field-level accuracy of 96.2 percent on the test set, up from the 94 to 96 percent baseline of manual entry. The monthly report cycle dropped from 14 working days to 36 hours: the pipeline ran overnight, the RAG assistant generated a draft narrative summary by 09:00 the next morning, and a finance analyst reviewed and approved the output by 12:00. The error rate on the final report fell to under 1 percent. The operations team gained access to supplier scorecards within 48 hours of month-end close, a 12-day improvement. The two analysts who previously spent 60 hours per cycle on this task were redeployed to supplier negotiation support. The n8n workflow was handed over with documentation, and the client’s own operations team could adjust routing rules without a developer. The RAG assistant was scoped to the indexed corpus only; it did not have internet access, and access was controlled at the Teams channel level.

    Lessons for Similar Teams

    • Fix the pilot scope before writing code. The eight-week timeline held because the audit deliverable defined exactly which invoices, which fields, and which ERP endpoints were in scope. Any new request during the pilot was treated as a change order with its own timeline, not a silent addition. Teams that skip this step routinely blow past their deadline by two to three weeks.
    • Self-host the orchestration layer when procurement asks where data lives. n8n on the client’s own infrastructure answered that question in one sentence. A managed SaaS orchestrator would have required a data processing agreement and a security review that added three to four weeks to the timeline.
    • Partition the RAG index by department. The operations assistant could not query finance data, and vice versa. This was enforced at the vector store level, not just at the Teams channel level. Without partitioning, a user in logistics could have pulled supplier contract terms from the finance index.
    • Log every human approval with a timestamp and user ID. Even though no regulation mandated it, the audit trail became the first thing the CFO asked for in the post-pilot review. The log showed exactly who approved the report, when, and what the model had drafted before approval.
    • Model-agnostic from day one. The client swapped the RAG model from OpenAI to the open-weight model in week three when procurement raised a data-residency concern. The n8n workflow did not change; only the model endpoint did. That swap cost two hours of configuration, not a re-architecture.
  • On-Premise AI Lead Qualification for a Swiss Professional Services Firm

    The Problem: 52-Hour Response Gaps and 6-Hour Reporting Cycles

    A 51-200 person professional services firm in Switzerland faces a specific operational bottleneck: inbound inquiries arrive across time zones and channels, but the sales team works 09:00-17:00 CET, Monday through Friday. A lead that lands at 22:00 on a Thursday waits 52 hours for a first substantive reply. In B2B professional services, that gap is not a minor inconvenience; it is a measurable conversion loss. The firm’s CRM holds the pipeline data, its Notion workspace holds the methodology documents, pricing sheets, and case studies, and its monthly reporting cycle consumes roughly 6 analyst-hours per month assembling numbers that already exist in the CRM.

    The problem is not a lack of data. It is a lack of a system that reads the data, classifies the inquiry, drafts a response, and files the report without a human touching each step. The firm does not need a new CRM or a new helpdesk. It needs an intelligent layer that sits on top of the tools it already runs, operates around the clock, and keeps every data point inside its own infrastructure because Swiss data-protection expectations and GDPR Article 32 make off-premise processing of client and prospect data a compliance risk the firm is not willing to take.

    Mechanism: On-Premise RAG, Open-Weight LLM, and the CRM Integration Layer

    The architecture has three components: a retrieval-augmented generation (RAG) pipeline, a conversational agent, and a reporting module. All three run on the client’s own hardware.

    The RAG pipeline ingests documents from the firm’s Notion workspace via the Notion API (version 2022-06-28), which exposes pages and blocks as JSON. Documents are chunked at heading boundaries, embedded with a sentence-transformer model (e.g., all-MiniLM-L6-v2, 384-dimensional vectors), and stored in a local Qdrant instance. At query time, the agent retrieves the top-5 chunks, builds a prompt with the retrieved context, and calls an open-weight LLM—Llama 3 70B or Mistral 8x7B—running on the firm’s GPU server. No document content or query text leaves the building.

    The conversational agent classifies each inbound inquiry into tiers: high-intent, mid-intent, low-intent. High-intent leads are routed to the CRM via its REST API with a structured summary. A human reviews every high-intent classification before the CRM record is created. The reporting module ingests CRM pipeline data and the firm’s Notion templates, drafts a structured monthly report with variance analysis, and queues it for human approval.

    The model-agnostic design means the firm can swap the LLM backend if a newer open-weight model outperforms the current one, without changing the RAG pipeline or the CRM integration.

    Trade-offs: Model Quality, Human Oversight, and Timeline

    The first trade-off is model quality versus data residency. A frontier API model (GPT-4o, Claude 3.5 Sonnet) would produce more nuanced lead classifications and better report narratives. But sending prospect names, firm details, and inquiry text to a third-party API violates the firm’s data-residency policy and complicates the GDPR Article 28 processor assessment. The open-weight model on-premise trades roughly 10-15% in classification accuracy for full data control. For a 51-200 person firm where the sales team reviews every high-intent lead anyway, that accuracy gap is acceptable.

    The second trade-off is the human-in-the-loop gate. Every high-intent classification requires a human approval before the CRM record is created. This adds roughly 90 seconds per lead and means the agent cannot fully automate the pipeline. But it eliminates the risk of a misqualified lead consuming a senior consultant’s time, and it satisfies the firm’s internal governance requirement that no AI output touches the sales pipeline without human sign-off.

    The third trade-off is the 4-week timeline. A full production rollout with monitoring, alerting, and a second channel would take 8-10 weeks. The 4-week pilot scopes to one workflow—lead qualification from inbound inquiries—and ships with a measured before/after baseline on cycle time and error rate. The firm accepts a narrower scope in exchange for a faster proof of value.

    Recommendation: Scope the Pilot to One Workflow, Measure the Delta

    The pilot targets lead qualification from inbound inquiries. The process audit in Week 1 maps the current workflow: inquiries arrive via email, web form, and phone, a sales associate manually classifies each one, drafts a first response, and logs the lead in the CRM. The baseline measurement captures cycle time (median 38 hours from inquiry to first response) and error rate (12% of leads misclassified in the prior quarter).

    Week 2 builds the RAG pipeline and connects the Notion API. Week 3 runs the agent in shadow mode against 200 historical inquiries, comparing its classifications to the human baseline. Week 4 adds the approval gate, connects the CRM write path, and measures the after-state. The target: reduce median first-response time to under 15 minutes for round-the-clock inquiries, and reduce misclassification rate to under 5%.

    The monthly reporting module ships in the same pilot. It ingests CRM pipeline data and the firm’s Notion reporting templates, drafts the monthly report, and queues it for analyst review. The target: reduce assembly time from 6 hours to 45 minutes of review and editing.

    The dedicated AI team operates as an embedded unit. The firm’s engineers and operations staff work alongside the team daily, not through a ticketing queue. This matters for a 4-week timeline: the team needs direct access to the Notion workspace, the CRM API credentials, and the firm’s GPU server, and it needs the operations staff available for the shadow-mode testing in Week 3.