Category: E-commerce and Retail

  • UAE E-commerce: LangGraph Document Extraction and Knowledge Search in Six Months

    The Problem: Routine Work That Should Not Require a Senior Headcount

    A 501-to-2,000-person e-commerce company in the UAE typically runs on a patchwork of Confluence pages, Notion databases, and a CRM that nobody has migrated in three years. The legal and compliance team spends roughly 30 percent of its week pulling product certificates, supplier contracts, and customs declarations out of PDFs, re-keying the data into spreadsheets, and answering the same “where is the compliance file for SKU 4471” question from the operations team. The problem is not a lack of tools; it is that the tools do not talk to each other, and the people who know where things live are the same people who are supposed to be reviewing contracts.

    The fix is not a new platform. It is a fixed-scope integration sprint that inserts an AI layer into the systems you already run. The sprint has a locked scope: one document type, one knowledge-search channel, one measured baseline. It does not replace your CRM, your ERP, or your helpdesk. It plugs into their APIs and adds a retrieval-augmented assistant on top. The architecture is model-agnostic: OpenAI or Anthropic APIs where speed matters, open-weight models on your own hardware where regulated data cannot leave the building. That last point is not optional in the UAE, where data-residency expectations under ISO 27001 Annex A.8.15 and the UAE Data Protection Law mean that a vendor-hosted model is a compliance risk, not just a cost line.

    The Audit: Picking the Workflow That Actually Moves the Needle

    The first two weeks of the engagement are the process audit. The team maps every document that enters the system: supplier invoices, customs declarations, product compliance certificates, internal policy PDFs, and the Confluence pages that hold the answers to “who approved this SKU for the Dubai market?” For each document type, the audit logs the current cycle time, the error rate, and the person who handles it. This is the before/after baseline that the pilot will be measured against.

    The audit also identifies which workflows are worth automating. Not everything is. A document type that appears four times a month and takes eleven minutes to process is not a pilot candidate. The target is a workflow that appears at least 200 times a month, has a measurable error rate above 2 percent, and touches a team that is already at capacity. In a typical UAE e-commerce operation, that is the supplier invoice and the product compliance certificate. The audit output is a one-page scope document that locks the pilot: one document type, one knowledge-search channel, one integration point.

    The scope is fixed. If the team discovers during the build that a second document type would be useful, that is a change request, not a scope expansion. This discipline is what separates an integration sprint from an open-ended consulting engagement, and it is what makes the six-month timeline credible.

    The Build: LangGraph Pipeline with a Human Approval Gate

    The pipeline is built on LangChain for the prompt and tool layer, and LangGraph for the stateful workflow. LangGraph matters here because the document extraction process is not a single call; it is a loop. The model extracts fields from the PDF, a confidence score is computed, and if the score is below 0.85 the item is routed to a human review queue. The human approves, corrects, or rejects. The corrected output is fed back into the training set. LangGraph models this loop as a graph with explicit nodes and edges, so the approval gate is a first-class part of the architecture, not a callback buried in a Python function.

    The knowledge-search assistant uses the same stack. Confluence and Notion both expose REST APIs that return page content as Markdown. The pipeline ingests that content, chunks it by heading, and indexes it in a vector store with metadata: page owner, last-updated date, access level. The LangGraph retrieval node queries the vector store, ranks the top five chunks, and passes them to the LLM for a grounded answer. The answer includes a citation to the source page and a confidence score. For legal and compliance queries, the output is routed to a human reviewer before it reaches the requester. This is the human-in-the-loop default: the model drafts, a person approves anything that touches a contract, a regulation, or a health-data reference.

    The model choice is deferred until the pipeline is working. Weeks two and three use an OpenAI or Anthropic API for speed. Weeks four and five swap to an open-weight model like Llama 3 70B on the client’s own hardware in a UAE data center. The LangGraph interface abstracts the model call, so the swap is a configuration change, not a rewrite.

    The Pilot: Six Weeks, One Document Type, One Measured Baseline

    The pilot runs for six to eight weeks. Week one is the audit and baseline. Weeks two through four are the build: the LangGraph pipeline, the Confluence and Notion API integration, the vector store, and the human review queue. Weeks five through six are the tuning cycle: the team watches the exception rate, adjusts the confidence threshold, and refines the prompt for the document types that are failing. The final two weeks are the measurement: the team compares the pilot’s cycle time and error rate against the baseline from the audit.

    The measurement is not a vanity metric. It is the document that goes to the CFO and the ISO 27001 auditor. The baseline report shows: before the pilot, the supplier invoice took 14 minutes to process and had a 4.2 percent error rate. After the pilot, it takes 3 minutes and the error rate is 0.8 percent. The knowledge-search assistant answered 78 percent of internal queries without a human, and the remaining 22 percent were routed to the review queue with a citation and a confidence score.

    The rollout decision is made at the end of week eight. If the error rate is below 1 percent and the cycle time is below 5 minutes, the pilot graduates to production. The production deployment adds monitoring: the exception rate becomes a KPI in the ISO 27001 operational monitoring plan, and any spike above 3 percent triggers a review of the model or the document format. The managed operation retainer covers the monitoring, the model updates, and the quarterly re-audit of the document types.

    Rollout and Managed Operation: What Happens After the Pilot

    The six-month timeline is not a single sprint. It is a sequence: the audit and pilot in months one and two, the rollout in month three, and the managed operation in months four through six. The rollout is not a big-bang deployment. It is a phased expansion: the first document type goes to production in week nine, the second in week eleven, and the knowledge-search assistant opens to the full team in week thirteen. Each phase has its own baseline measurement and its own exception-rate threshold.

    The managed operation phase is where the engagement stops being a project and starts being a service. The vendor monitors the exception rate, the model performance, and the integration health. If the Confluence API changes its response format, the vendor patches the ingestion layer within 48 hours. If the document format shifts because a new supplier starts sending a different invoice layout, the vendor re-trunes the extraction prompt and re-runs the baseline. The client’s team does not need to hire a data scientist or an ML engineer to keep the system running. That is the point of scaling operations without new hires: the AI layer absorbs the routine work, and the human team focuses on the exceptions and the decisions that actually require judgment.

    The ISO 27001 audit trail is maintained throughout. Every model call, every human approval, every exception routing is logged with a timestamp, the user ID, and the document reference. The logs are stored in the client’s own infrastructure, not in a vendor’s cloud. This is the difference between a system that passes an audit and a system that is built to be audited.

  • AI Workflow Automation vs. Customer Response for E-Commerce in the UAE

    What Is Being Compared

    The two options under comparison are distinct in function, even though both use the same underlying model layer. AI workflow automation targets internal back-office processes: invoice processing, document extraction, and data entry. The goal is to reduce cycle time and error rate in operations and supply chain. Round-the-clock customer response targets external-facing channels: ticket triage, first-response agents, and voice. The goal is to cut first-response time and maintain service levels across time zones. Both options use the OpenAI API as the model layer, integrate with existing tools via API, and ship with a human-in-the-loop approval step. The difference is the workflow being automated and the metric that defines success.

    Criteria for Comparison

    We judge each option against seven criteria that matter to an 11–50 person e-commerce team in the UAE with no specific compliance constraints:

    • Cycle time reduction (internal workflow) vs. first-response time (customer-facing)
    • Error rate (data entry, invoice matching) vs. escalation rate (ticket misclassification)
    • Integration complexity with existing ERP, helpdesk, and documentation tools
    • Human-in-the-loop overhead (approval steps per transaction)
    • Cost per transaction (API call volume, token usage)
    • Time to value within the 8-week fixed-scope pilot
    • Scalability beyond the pilot scope (additional workflows or channels)

    Comparison Table

    Criterion AI Workflow Automation (Invoice Processing) Round-the-Clock Customer Response
    Primary metric Cycle time (hours per invoice) First-response time (minutes per ticket)
    Error rate target <2% mismatch or misclassification <5% misrouted or escalated tickets
    Integration points ERP, accounting software, Notion/Confluence for audit trail Helpdesk, messaging platform, CRM
    Human-in-the-loop Approval before payment or data entry Approval for high-value or sensitive tickets
    API call volume Moderate (one call per invoice) High (one call per ticket, 24/7)
    Time to value in 8 weeks Measurable by week 6 Measurable by week 4
    Scalability Add more invoice types or suppliers Add more channels or languages

    When Each Option Wins

    AI workflow automation wins when the team’s bottleneck is internal: invoice processing is slow, error-prone, and consumes operator time that could go to supply-chain planning. For a 15-person e-commerce team, reducing invoice cycle time from 4 hours to 30 minutes frees up roughly 3.5 operator-hours per invoice. Over 200 invoices per month, that is 700 hours—enough to hire one additional operations analyst or reduce overtime. The fixed-scope pilot delivers a clear before/after baseline on cycle time and error rate, making the business case straightforward.

    Round-the-clock customer response wins when the team’s bottleneck is external: first-response time is high, tickets are piling up, and the team cannot cover all time zones. For an e-commerce business in the UAE serving customers across the Gulf and beyond, a 24/7 AI first-response agent can cut first-response time from 4 hours to 15 minutes. The pilot measures escalation rate and customer satisfaction, and the human-in-the-loop step ensures that high-value or sensitive tickets are routed to a person.

    Recommendation

    For an 11–50 person e-commerce team in the UAE with no specific compliance constraints, AI workflow automation for invoice processing is the stronger first pilot. The reasons are concrete: the workflow is high-volume and repetitive, the success metric (cycle time) is easy to measure, and the human-in-the-loop approval step (before payment) reduces risk. The 8-week timeline is sufficient to audit the process, integrate with the ERP and Notion or Confluence for the audit trail, and deliver a before/after baseline. The OpenAI API is appropriate for the quality of document extraction and classification required. If the pilot meets the target—say, cycle time reduced by 70% and error rate below 2%—the team can scale to additional workflows or add customer-facing automation in a second pilot.

  • Deploying a RAG Assistant for Order Status Updates in Swiss E-Commerce

    The Problem: Senior Support Staff Buried in Routine Order Status Tickets

    Your support team at a 2,000+ employee e-commerce company in Switzerland handles thousands of order and shipment status inquiries weekly. Senior agents spend 40-60% of their time on routine lookups: “Where is my package?” “Why is my order delayed?” This work does not require judgment, but it consumes the people who should be handling complex escalations, refund disputes, and customer retention conversations. The EU AI Act, which applies to Swiss companies serving EU customers, requires transparency when AI systems interact with users. You need a retrieval-augmented knowledge assistant that drafts accurate responses from your order-management system and shipping carrier data, integrates with Zendesk or Intercom, and keeps a human in the loop for anything touching refunds or contract terms. The goal: cut first-response time from hours to minutes, reduce error rate on shipping information, and free senior staff for high-value work within a 3-month integration sprint.

    Prerequisites: What You Need Before the Sprint Starts

    Before starting the integration sprint, confirm these are in place:

    • Zendesk or Intercom API access: OAuth 2.0 tokens with read/write permissions for tickets, macros, and webhooks. Test with a sandbox account first.
    • Order-management system (OMS) API: Read access to order status, tracking numbers, and shipping carrier data. If you use Shopify, SAP Commerce, or a custom OMS, document the endpoint schema.
    • Shipping carrier APIs: Integration with at least your top two carriers (e.g., Swiss Post, DHL) for real-time tracking events.
    • PostgreSQL 15+ with pgvector extension: CREATE EXTENSION vector; Run on a dedicated instance with at least 16 GB RAM for 500k+ vectors.
    • LLM endpoint: OpenAI API key (gpt-4o or claude-3-5-sonnet) for drafting, or an on-prem Llama 3 70B instance if customer PII cannot leave your infrastructure.
    • EU AI Act compliance documentation: A data-protection impact assessment (GDPR Article 35) and a model card for each LLM endpoint.
    • Baseline metrics: Export 30 days of ticket data from Zendesk/Intercom. Calculate average first-response time, resolution rate, and error rate on shipping-related tickets.

    Step 1: Audit the Workflow and Establish a Baseline

    Run a process audit on your last 90 days of support tickets. Filter for order and shipment status inquiries: “Where is my order?” “Tracking number not working” “Delivery delayed.” Count the volume, measure average handling time, and identify the top five questions. For a 2,000+ employee e-commerce company, this typically represents 35-50% of total ticket volume. Export the data to a CSV with columns: ticket_id, subject, category, first_response_time, resolution_time, agent_id, error_flag. Calculate the baseline: if your average first-response time is 4 hours and error rate on shipping information is 8%, those are your targets to beat. Document this baseline in a one-page report. This becomes the measurement framework for the pilot and rollout phases.

    Step 2: Build the RAG Pipeline with pgvector

    Build the retrieval layer using pgvector. Chunk your knowledge base: shipping policies, carrier SLAs, return procedures, and order status definitions. Use a 512-token chunk size with 50-token overlap. Generate embeddings with OpenAI text-embedding-3-small (1536 dimensions) or bge-base-en-v1.5 (768 dimensions) if you prefer open-weight models. Load into PostgreSQL:

    CREATE TABLE documents (
      id SERIAL PRIMARY KEY,
      content TEXT,
      metadata JSONB,
      embedding vector(1536)
    );
    CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
    

    Set ef_search = 64 for sub-10 ms recall. Test with 20 sample queries: “Where is my order with tracking number XYZ?” Verify that the top-5 retrieved chunks contain the relevant shipping policy and carrier SLA. If recall is below 90%, adjust chunk size or add metadata filters (e.g., WHERE metadata->>'carrier' = 'DHL').

    Step 3: Integrate with Zendesk or Intercom via Webhooks

    Connect the RAG pipeline to Zendesk or Intercom. For Zendesk: create a webhook on ticket creation that triggers your RAG service. The service retrieves relevant chunks, calls the LLM endpoint with a system prompt: “You are a support assistant for [Company]. Use only the retrieved context to draft a response. If the context does not contain the answer, say so. Do not invent tracking numbers or delivery dates.” Post the drafted response to the ticket via the API with a RAG-drafted tag. For Intercom: use the Events API to trigger on ticket.created and the Agent Inbox API to post the draft. Store the correlation ID (ticket_id + timestamp) in a log table for audit trails. This satisfies EU AI Act Article 50 transparency requirements: users are informed they are interacting with an AI, and every response is traceable to its source documents.

    Step 4: Add Human-in-the-Loop Approval for Sensitive Actions

    Implement the human-in-the-loop approval workflow. Any RAG-drafted response that touches refunds, address changes, or contract terms must be approved by a human before sending. In Zendesk, create a custom field ai_approval_status with values: pending, approved, rejected. When the RAG service posts a draft, set ai_approval_status = pending and assign the ticket to a supervisor queue. The supervisor reviews the draft, the retrieved context, and the LLM’s confidence score. If approved, the ticket moves to approved and the response sends. If rejected, the supervisor edits or reassigns. Log every approval decision with the supervisor’s user ID and timestamp. This workflow is mandatory under EU AI Act Article 50 for any AI system that makes decisions affecting consumers. For a 3-month sprint, build a simple approval UI in React or use Zendesk’s built-in ticket views filtered by ai_approval_status = pending.

    Step 5: Pilot with 10-20% of Tickets and Measure

    Run the pilot with 10-20% of order-status tickets for two weeks. Route a subset of tickets (e.g., all tickets tagged order_status from a specific region or carrier) to the RAG assistant. Measure: first-response time (target: under 15 minutes vs. baseline 4 hours), resolution rate (target: 80%+ first-contact resolution), and error rate on shipping information (target: under 2% vs. baseline 8%). Sample 5% of AI-drafted responses weekly. Compare each against the OMS and carrier API data. If the assistant states a delivery date, verify it matches the carrier’s tracking event. If error rate exceeds 2%, pause the pilot, re-index the knowledge base, and adjust the LLM prompt to require citation of specific tracking events. Document every error in a log with the ticket ID, the incorrect claim, and the correct data from the OMS. This log feeds into the EU AI Act model card and the GDPR Article 35 impact assessment.

  • Automating the Monthly Compliance Report at a 201-500-Person UAE E-Commerce Firm

    The Monthly Report That Eats Fourteen Hours

    The monthly compliance report at a 201-500-person e-commerce firm in the UAE is not a single task. It is a chain of twelve to eighteen manual steps: pulling sales figures from the ERP, reconciling returns from the helpdesk, extracting vendor payment data from the accounting system, formatting the narrative summary, and filing the result with the internal compliance officer. The person who owns this workflow — usually a senior operations analyst or a compliance coordinator — spends 12 to 16 hours per cycle, and the error rate on manual transcription sits between 3 and 7 percent. A single mis-keyed figure can trigger a late filing or a wrong vendor payment, and the cost of a correction is not just the hours to fix it but the reputational friction with the internal audit team.

    The pain is structural, not personal. The analyst is not slow; the data is scattered across four systems that do not talk to each other. The ERP exposes a REST API, but the helpdesk only offers a CSV export. The vendor payment data lives in a spreadsheet that a finance clerk updates by hand. The analyst is, in effect, a human ETL pipeline, and the monthly deadline makes the work feel urgent even though the underlying process has not changed in three years.

    Why RPA and Vendor Reports Do Not Fix This

    The first common response is to buy a RPA tool — UiPath, Automation Anywhere, or a lighter-weight option — and have a consultant build a bot that clicks through the ERP, the helpdesk, and the spreadsheet. RPA works when the screens are stable and the data is in a predictable location. In a 201-500-person e-commerce firm, the screens are not stable. The ERP vendor ships a quarterly UI update. The helpdesk CSV export changes column order when the vendor upgrades. The spreadsheet has a new tab every month because the finance clerk “reorganized” it. The RPA bot breaks, and the consultant is no longer on retainer. The analyst goes back to manual work, now with a broken bot to ignore.

    The second common response is to ask the ERP or helpdesk vendor to build a custom report. This takes six to ten weeks of vendor project time, costs EUR 15 000 to EUR 40 000, and delivers a static PDF that still requires a human to interpret and file. The vendor has no incentive to build a report that spans three of its own products plus a spreadsheet. The result is a report that is accurate but slow, and the analyst still spends four to six hours on interpretation and formatting.

    The third response is to hire another analyst. This doubles the headcount cost without fixing the root cause: the data is still scattered, the process is still manual, and the new analyst inherits the same 14-hour cycle. The firm has bought time, not capacity.

    A Fixed-Scope Pilot on the Claude API

    The path that works for a firm at this stage — no AI in production yet, a 3-month timeline, a fixed-scope pilot — is a workflow-orchestration layer that sits on top of the existing systems rather than replacing them. The architecture is model-agnostic, but for a monthly compliance report where the narrative summary and the exception flagging benefit from strong language understanding, the Anthropic Claude API is the right fit. The system pulls data from the ERP via its REST API, triggers on a webhook from the helpdesk when a new returns batch lands, and reads the vendor payment spreadsheet through a lightweight parser. The Claude API handles the classification of exceptions, the drafting of the narrative summary, and the flagging of any figure that deviates from the prior month by more than a set threshold.

    The human-in-the-loop step is non-negotiable. The model drafts the report; a named compliance officer reviews it, corrects any flagged fields, and signs off. The approval log is stored as part of the audit trail. The system does not file the report automatically. It prepares it, flags it, and waits for the human. This keeps the cycle time low while ensuring that no number reaches the internal audit team without a person having seen it.

    The pilot ships with a measured before/after baseline: cycle time, error rate, and the number of manual steps. The target is to cut the 14-hour cycle to under 2 hours and reduce transcription errors to zero. The scope is locked in writing before development starts.

    From Pilot to Internal Knowledge Search

    The pilot is not the end of the story. The same orchestration layer that automates the monthly report can be extended to the internal knowledge search use case. The firm’s SOPs, vendor contracts, past compliance filings, and CRM records are chunked, embedded, and stored in a vector database. When an analyst asks, “What was the return rate for Q3 in the Gulf region?” the system retrieves the relevant chunks, passes them to the Claude API as context, and generates a cited answer with a link to the source document. This is a retrieval-augmented generation layer, not a chatbot. The accuracy depends on the quality of the source documents, so the process audit includes a document-hygiene pass before the RAG layer is built.

    The integration is through custom REST APIs and webhooks, not through a new middleware platform. The ERP already exposes a REST API. The helpdesk already fires webhooks on new tickets. The vendor payment spreadsheet is read by a parser that runs on a schedule. No new infrastructure is required. The system plugs into what the firm already runs.

    The 3-month timeline is realistic if the source systems expose clean APIs. Month one: process audit, baseline measurement, architecture design. Month two: build and integration. Month three: testing, human-in-the-loop validation, and the before/after measurement. If the audit reveals that data is trapped in PDFs with no API, add two to four weeks for a data-extraction layer.

    Five Steps to Start in Month One

    The first step is a one-to-two-week process audit. The goal is not to design the solution but to measure the baseline: how many hours the current monthly report takes, how many manual steps, the error rate over the last three cycles, and which systems the data comes from. The audit produces a one-page scorecard ranking the workflows by volume, error cost, and data availability. The pilot picks the top-ranked workflow that also has a clean data path.

    The second step is to name a single owner for the workflow. This is the person who will approve the AI’s output, correct flagged fields, and sign off on the report. Without a named owner, the human-in-the-loop step becomes a group chat, and the cycle time does not improve.

    The third step is to confirm API access. The ERP vendor must grant read access to the relevant endpoints. The helpdesk must confirm that webhooks can be configured for the returns batch. The vendor payment spreadsheet must be stored in a location the parser can reach. If any of these are blocked, the timeline stretches, and the pilot scope must be adjusted.

    The fourth step is to lock the pilot scope in writing. The deliverable, the acceptance criteria, the deadline, and the before/after metrics are all specified before development starts. The client pays for a known outcome, not an open-ended retainer.

    The fifth step is to run the pilot and measure. The pilot ships the automation, the integration, and a one-page report comparing baseline to actual. If the numbers move, the firm scales the pattern to adjacent workflows. If they do not, the firm has the baseline data and a clear diagnosis of why.

  • Cutting Contract Review Cycle Time in Swiss E-Commerce: A 2-Week AI Pilot

    The Contract Review Bottleneck in Swiss E-Commerce

    A 201-500 employee e-commerce company in Switzerland runs its legal and compliance function on a small team. Contract review for vendor agreements, data processing agreements, and customer-facing terms consumes 4 to 6 hours per document. The legal team tracks cycle time manually in a spreadsheet, and error rate on standard clauses sits at 12 to 18 percent because reviewers work through queues without a consistent precedent library. First-response time on internal compliance queries from the sales and operations teams averages 2 to 3 business days because the legal team is buried in contract work. The cost per support ticket that touches a contract question runs 35 to 50 Swiss francs in legal time, and the team has no baseline to measure improvement. The company has run two isolated AI pilots in the last 18 months, neither of which reached production because the scope was undefined and the integration with existing systems was never planned.

    Why Isolated Pilots Stall in Legal and Compliance

    Most companies in this position reach for one of three approaches, and each fails in a predictable way. The first is a generic LLM wrapper: a legal team member pastes a contract into ChatGPT and asks for a summary. This produces plausible-sounding output that misses jurisdiction-specific clauses, Swiss data protection requirements under the revised nFADP, and the company’s own precedent language. The second is a RAG pipeline built on a single document store without a structured extraction layer. The retrieval step finds relevant clauses, but the extraction step that pulls out party names, payment terms, and liability caps is brittle and requires manual correction on 30 to 40 percent of documents. The third is a full vendor platform that replaces the existing CRM and document management system. The integration cost alone exceeds the annual legal budget for a 201-500 employee firm, and the migration timeline stretches past 12 months. None of these approaches ship a measured before/after baseline, so the company cannot prove the pilot reduced cycle time or error rate.

    A Fixed-Scope Pilot That Ships in Two Weeks

    The fix starts with a 2-week AI automation audit that maps the contract review workflow end to end. Forfis interviews the legal team, identifies the top 3 to 5 document types by volume and error rate, and scores each on automation feasibility and data sensitivity. The audit delivers a fixed-scope pilot proposal on the single workflow with the best risk-to-reward ratio, typically standard vendor contracts. The pilot architecture uses a model-agnostic stack: OpenAI or Anthropic APIs for classification and drafting where quality matters, open-weight models on the client’s own hardware where regulated data cannot leave the building. A pgvector embeddings search layer indexes the company’s contract templates, precedent clauses, and compliance checklists from Notion or Confluence, so the AI agent retrieves relevant language before drafting. The system plugs into the existing CRM and helpdesk through their APIs rather than replacing them. Every pilot ships with a measured before/after baseline on cycle time and error rate, and the human-in-the-loop approval step ensures no contract touches a counterparty without legal sign-off.

    How to Start: Five Concrete Steps

    Week 1 of the audit: Forfis maps the current contract review process, identifies the top 3 to 5 document types by volume, and records baseline cycle time and error rate on a sample of 50 to 100 historical contracts. The team interviews the legal and compliance staff to understand which clauses are non-negotiable and which can be auto-classified. Week 2: the team builds a proof-of-concept extraction pipeline on the sample, measures the before/after delta, and delivers a fixed-scope pilot proposal with cost, timeline, and EU AI Act compliance controls. The pilot itself runs 4 to 6 weeks and ships with a measured baseline. From there, rollout extends to additional document types and the managed operation phase handles model updates, drift monitoring, and compliance reporting. The first step is to schedule the audit. The second is to gather 50 to 100 historical contracts in a shared Notion or Confluence workspace. The third is to identify the single workflow with the highest volume and error rate. The fourth is to define the success metric: cycle time reduction and error rate drop. The fifth is to assign a legal owner who will approve every AI-drafted output during the pilot.

  • Swiss E-commerce Retailer Cuts Reporting Cycle Time 70% with AI Automation

    Background and Challenge

    This case study is a composite based on patterns observed in the field. It does not represent a single named customer but reflects common challenges and solutions in the e-commerce and retail sector in Switzerland.

    Background
    A mid-sized Swiss e-commerce retailer with 1,200 employees operates across DACH markets. The company uses a custom-built CRM and ERP system, with data stored in on-premise servers. The sales team of 45 handles lead qualification and monthly reporting manually, using spreadsheets and email. The company has no AI in production yet and is looking to reduce manual back-office work while improving lead qualification accuracy.

    Challenge
    The sales team spends 12 hours per week on monthly reporting, manually aggregating data from the CRM, ERP, and web analytics. The process is error-prone, with a 10% error rate in data entry. Lead qualification is inconsistent, with 30% of leads being misclassified, leading to lost opportunities. The company faces GDPR compliance requirements and a deadline to implement improvements before the Q4 peak season.

    Approach
    Forfis conducted an AI automation audit, identifying monthly reporting and lead qualification as high-impact use cases. A fixed-scope pilot was designed to automate these workflows using LangChain and LangGraph for workflow orchestration. The system integrates with the existing CRM and ERP via custom REST APIs and webhooks. A human-in-the-loop model ensures that AI-generated reports and lead scores are reviewed by a human before finalization. The pilot was deployed in two weeks, with a measured before/after baseline on cycle time and error rate.

    Outcome
    The pilot reduced monthly reporting cycle time from 5 days to 1 day, a 70% improvement. The error rate decreased from 10% to 2%, an 80% reduction. Lead qualification accuracy improved from 70% to 95%, with a 25% increase in qualified leads passed to sales. The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building.

    Lessons

    • Start with a fixed-scope pilot to demonstrate ROI quickly.
    • Use a human-in-the-loop model to ensure accuracy and compliance.
    • Integrate with existing systems via APIs rather than replacing them.
    • Measure before/after baselines to quantify impact.
    • Choose a model-agnostic architecture to future-proof the solution.

    Approach: AI Automation Audit and Pilot Design

    The AI automation audit identified two high-impact use cases: monthly reporting and lead qualification. The audit mapped existing workflows, identified bottlenecks, and evaluated the feasibility of automating specific tasks. The results were a prioritized list of use cases, with estimated ROI and implementation complexity.

    Monthly Reporting
    The current process involves manually aggregating data from the CRM, ERP, and web analytics. The sales team spends 12 hours per week on this task, with a 10% error rate in data entry. The AI system automates data collection, validation, and report generation. It uses LangChain to chain prompts and tools, and LangGraph to define stateful, multi-step workflows. The system integrates with the existing CRM and ERP via custom REST APIs and webhooks, ensuring data integrity and real-time updates.

    Lead Qualification
    The current process is inconsistent, with 30% of leads being misclassified. The AI system uses a classification model to score leads based on predefined criteria, such as company size, industry, and engagement level. The model is trained on historical data and fine-tuned using feedback from the sales team. A human-in-the-loop model ensures that AI-scored leads are reviewed by a human before they are passed to sales, ensuring accuracy and context.

    GDPR Compliance
    The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building. Data minimization is implemented, and data subjects can exercise their rights. The legal basis for processing is documented, and third-party AI APIs are GDPR-compliant. The system uses open-weight models on the client’s own hardware, ensuring that regulated data does not leave the building.

    Outcome: Measured Impact on Cycle Time and Error Rate

    The pilot was deployed in two weeks, with a measured before/after baseline on cycle time and error rate. The system was integrated with the existing CRM and ERP via custom REST APIs and webhooks, ensuring seamless data flow. The human-in-the-loop model was implemented, with a review dashboard for the sales team to approve AI-generated reports and lead scores.

    Cycle Time
    The monthly reporting cycle time was reduced from 5 days to 1 day, a 70% improvement. The AI system automates data collection, validation, and report generation, eliminating manual data entry and aggregation. The sales team spends 2 hours per week on review and approval, compared to 12 hours previously.

    Error Rate
    The error rate in monthly reporting decreased from 10% to 2%, an 80% reduction. The AI system validates data in real-time, flagging anomalies and inconsistencies. The human-in-the-loop model ensures that errors are caught and corrected before the report is finalized.

    Lead Qualification Accuracy
    Lead qualification accuracy improved from 70% to 95%, with a 25% increase in qualified leads passed to sales. The AI system scores leads based on predefined criteria, and the human-in-the-loop model ensures that misclassified leads are corrected. The sales team reports a 15% increase in conversion rates, attributed to more accurate lead qualification.

    GDPR Compliance
    The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building. The legal basis for processing is documented, and data subjects can exercise their rights. The system uses open-weight models on the client’s own hardware, ensuring that regulated data does not leave the building.

    Lessons for Similar Teams

    The pilot demonstrated significant improvements in cycle time, error rate, and lead qualification accuracy. The system is GDPR-compliant and integrated with existing systems via APIs. The human-in-the-loop model ensures accuracy and compliance, while the model-agnostic architecture provides flexibility and future-proofing.

    Scalability
    The system can be scaled to automate other workflows, such as invoice processing and document extraction. The model-agnostic architecture allows for switching between different AI models, based on cost, performance, and compliance requirements. The system can be extended to other departments, such as marketing and customer service, with minimal changes.

    Cost Efficiency
    The pilot reduced manual back-office work by 80%, saving 10 hours per week. The cost of the AI system is offset by the reduction in manual effort and the increase in qualified leads. The system is cost-effective, with a payback period of less than 3 months.

    Risk Mitigation
    The human-in-the-loop model mitigates the risk of errors and ensures compliance with regulations. The model-agnostic architecture mitigates vendor lock-in and allows for future-proofing. The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building.

    Next Steps
    The company plans to roll out the system to other departments, such as marketing and customer service. The system will be extended to automate other workflows, such as invoice processing and document extraction. The company will continue to measure the impact of the system on key metrics, such as cycle time, error rate, and lead qualification accuracy.

  • 8-Week AI Automation Pilot for Lead Qualification in Austrian E-Commerce

    1. Verify the process audit scope and baseline metrics

    The audit is not a generic AI strategy session. It is a targeted assessment of the lead qualification workflow, from first touch to sales handoff. You map every step, identify where errors occur, and measure the current cycle time. The output is a prioritized list of automation opportunities, ranked by error rate and business impact. For a 51-200 employee e-commerce firm, this typically means 3 to 5 workflows, with lead qualification as the most common first candidate. The audit should take 1 to 2 weeks and produce a one-page roadmap with a clear recommendation on which workflow to automate first. This is the foundation for the entire 8-week engagement, and skipping it leads to wasted effort on the wrong process.

    2. Configure the human-in-the-loop approval gate

    The pilot must run on a single workflow, not multiple. For lead qualification, this means the AI classifies incoming leads, extracts key data, and drafts a response, but a human approves every action before it is sent. The human-in-the-loop gate is not optional; it is a compliance requirement under ISO 27001 and a practical safeguard against model errors. You define the approval rules in Notion or Confluence, so every decision is documented and auditable. The pilot should process at least 200 to 500 leads to generate statistically meaningful data. If your lead volume is lower, extend the pilot to 8 weeks to capture sufficient volume. The goal is to measure a reduction in error rate and cycle time, not to achieve 100% automation.

    3. Deploy open-weight models on-premise for regulated data

    For regulated data, open-weight models on your own hardware are the right choice. Llama 3 or Mistral can run on a single GPU server, ensuring no data leaves your infrastructure. This is critical for ISO 27001 compliance and for handling customer data under GDPR. The trade-off is that open-weight models may have lower quality on complex reasoning tasks, but for lead qualification, which is largely classification and extraction, they perform well. You can use a hybrid approach: open-weight for data processing and classification, and a commercial API for any free-text summarization that requires higher quality. The model must be versioned, and every prompt and output must be logged for audit purposes.

    4. Integrate with Notion or Confluence for documentation and audit trails

    The AI system must integrate with your existing CRM, helpdesk, and knowledge base. For this scenario, Notion or Confluence is the knowledge base, and the integration is via API. The AI system reads the process documentation, model prompts, and approval rules from Notion, and writes the results back. This ensures that the workflow is transparent and auditable. The integration should be tested in the first week of the pilot, before any leads are processed. If the integration fails, the entire pilot is compromised. You need a clear data flow diagram that shows how data moves from the lead source, through the AI system, to the CRM, and back to Notion for documentation.

    5. Document the ISO 27001 compliance controls for the AI system

    ISO 27001 requires you to document the information security controls for any system that processes sensitive data. For an AI workflow, this means documenting the data flow, access controls, model versioning, and human approval gates. You must show that the AI system is subject to the same security controls as your other business systems. Specifically, you need to document how the model is trained or fine-tuned, how prompts are managed, how outputs are validated, and how incidents are handled. The audit trail for every automated decision must be retrievable and reviewable. This documentation is not a one-time task; it must be updated as the workflow evolves.

    6. Measure the before-and-after baseline for cycle time and error rate

    The pilot should run for 4 to 6 weeks, with the first 1 to 2 weeks dedicated to integration and data mapping. You need enough volume to measure a statistically meaningful difference in error rate and cycle time. For lead qualification, that means processing at least 200 to 500 leads through the automated workflow and comparing the results against the manual baseline. If your lead volume is lower, extend the pilot to 8 weeks to capture sufficient data. The remaining 2 to 4 weeks of the 8-week timeline are for refinement, human-in-the-loop tuning, and documentation. The goal is a measurable reduction in both cycle time and error rate, with the error rate reduction being the primary KPI for this engagement.

    7. Identify and mitigate the top 5 pitfalls in the 8-week timeline

    The most common pitfalls are: 1) Automating the wrong process, which wastes the 8-week timeline. 2) Skipping the baseline measurement, which makes it impossible to prove ROI. 3) Not defining clear human approval gates, which creates compliance risk. 4) Over-relying on the AI without sufficient human review, which leads to errors in regulated data. 5) Failing to document the workflow in Notion or Confluence, which breaks ISO 27001 audit trails. 6) Choosing a model that is too complex for the task, which increases cost and latency without improving accuracy. Each of these can be avoided with proper scoping and governance. The 8-week timeline is tight, so every week must be planned and executed with precision.

  • Swiss E-commerce Cuts Invoice Processing to 3 Hours with On-Premise AI

    Background: A Swiss E-commerce Operator at 300 Headcount

    This case study is a composite based on patterns observed across Forfis engagements. We do not name real customers. The company described here is a mid-sized Swiss e-commerce operator with roughly 300 employees, running a multi-channel retail operation across DACH and Western Europe. The stack is a mix of a legacy ERP for inventory and finance, a modern CRM for customer relationships, and a helpdesk platform for internal and supplier communications. The operations team handles 1,200 to 1,800 supplier invoices per month, plus a monthly consolidated report that feeds into the finance close. The company is in the AI-native operations stage: leadership has approved AI investment, but the team has not yet built internal capability to deploy and maintain AI workflows. The engagement ran over 8 weeks, delivered by a dedicated Forfis AI team embedded with the client’s operations group.

    Challenge: 12 Hours a Week of Manual Invoice Entry and a Fixed Monthly Close

    The operations team spent an estimated 12 to 15 hours per week on manual invoice processing: extracting line items from PDFs, matching them against purchase orders in the ERP, flagging discrepancies, and entering validated data. The monthly consolidated report required pulling data from three systems, reconciling it, and formatting it for the finance close. The error rate on the baseline was 4.2 percent on invoice line items, with a 3-day average cycle time from receipt to posting. The pressure was twofold: the monthly close deadline was fixed, and the team had lost two senior operators to attrition in the prior quarter. Leadership wanted to reduce manual back-office work without replacing the existing ERP or CRM, and without sending supplier or financial data to a third-party cloud. The compliance posture was internal: no regulatory mandate, but the finance director required that all financial data remain on-premise.

    Approach: On-Premise Open-Weight Models with a Human-in-the-Loop Approval Layer

    Forfis ran a two-week process audit to map the invoice workflow end-to-end and capture baseline metrics. The pilot scope was fixed: automate invoice extraction, PO matching, and discrepancy flagging, plus generate the monthly consolidated report from the same data pipeline. The architecture used open-weight models deployed on the client’s own hardware, so all invoice and financial data stayed on-premise. The AI layer connected to the ERP and helpdesk through custom REST APIs and webhooks: the ERP pushed new invoices via webhook, the AI service processed them, and validated records were written back through the ERP’s REST API. Discrepancies were pushed to the helpdesk as tickets for human review. The human-in-the-loop layer was built into the workflow: the model drafted and classified, a person approved anything touching a financial transaction. The dedicated Forfis team handled technical planning, product design, and full-cycle development over the 8-week timeline.

    Outcome: Cycle Time Down 75 Percent, Error Rate Under 1 Percent

    After the 8-week engagement, the measured results were: cycle time on invoice processing dropped from 12 to under 3 hours per week, a reduction of roughly 75 percent. The error rate on invoice line items fell from 4.2 percent to under 1 percent. The monthly consolidated report, which previously took 2 to 3 days of manual reconciliation, was generated automatically from the same data pipeline and required only a 30-minute human review. The human-in-the-loop approval queue handled roughly 8 to 12 percent of invoices that required manual review, down from 100 percent. The operations team redirected the freed capacity to supplier relationship management and exception handling. The finance director confirmed that all data remained on-premise throughout the pilot and rollout, and the monthly close process was unchanged in structure but faster in execution. The system is now in managed operation with Forfis monitoring model performance and handling drift.

    Lessons for Similar Teams

    • Baseline before you build. The 4.2 percent error rate and 12-hour cycle time were captured during the audit, not estimated. Without that baseline, the outcome metrics would be unverifiable. Any team automating a back-office workflow should measure the current state before touching the process.
    • One workflow, not five. The pilot scope was fixed to invoice processing and monthly reporting. Attempting to automate the entire back-office in 8 weeks would have diluted the team’s focus and made the baseline unmeasurable. Sequence the rollout: prove one workflow, then expand.
    • On-premise is not a constraint, it is a design choice. The open-weight model on the client’s hardware was not a compromise. It was the right fit for the data residency requirement, and the model-agnostic architecture meant the team could swap models without re-architecting the integration layer.
    • Human-in-the-loop is the default, not a fallback. The approval layer was built into the workflow from day one, not added after a failure. The 8 to 12 percent manual review rate is a feature, not a bug: it keeps the team in control of financial transactions while the AI handles the volume.
    • Integration through existing APIs, not replacement. The custom REST API and webhook layer connected to the ERP and helpdesk without requiring data migration. This kept the project within the 8-week timeline and avoided the risk of a parallel system.
  • RAG Assistant for Order Status: 2-Week Pilot in Austrian E-commerce

    The Problem: Manual Order Status Queries in a 25-Person E-commerce Team

    A 25-person e-commerce operation in Vienna handles 400-600 customer inquiries daily, most of them asking where their order is. The support team spends 3-4 hours per agent per day on these repetitive queries, pulling up order management screens, checking carrier tracking numbers, and drafting responses. First-response time averages 6 hours, and document turnaround for shipping confirmations takes 1-2 business days. The business function is customer support, but the bottleneck is manual data retrieval and response drafting, not the actual customer interaction. The need is clear: cut first-response time to under 2 minutes and reduce document turnaround to same-day processing, without adding headcount or replacing existing systems. The solution must work within PCI DSS constraints because the support team occasionally handles refund requests that touch cardholder data, and it must integrate with Google Workspace, which the team already uses for email and calendar management. The pilot scope is one specific workflow: order and shipment status updates, chosen because it is high-volume, rule-based, and has clear before/after metrics to measure success.

    Architecture: Open-Weight Models On-Premise for PCI DSS Compliance

    The architecture uses open-weight models running on the client’s own hardware, not cloud APIs. This is a deliberate choice driven by PCI DSS compliance: cardholder data and transaction details must not leave the client’s controlled infrastructure. The model is a 7B-parameter open-weight variant, fine-tuned on the client’s historical support tickets and order management documentation. It runs on a single GPU server in the client’s data center, with all inference happening locally. The retrieval layer connects to the client’s order management system and shipping carrier APIs via standard REST endpoints, pulling real-time order status, tracking numbers, and delivery windows for each query. The assistant does not store transaction data; it retrieves it on demand, which means the model never has persistent access to sensitive information. This architecture satisfies PCI DSS requirement 3.4, which mandates that cardholder data be rendered unreadable at rest, and requirement 4, which requires encryption of data in transit. The model-agnostic design means that if the client later wants to use a different model for a different workflow, the retrieval layer and integration code remain unchanged.

    Pilot Scope: Two-Week Deployment on Order Status Queries

    The pilot runs for two weeks, starting with a process audit that maps the current workflow for order status queries. The audit identifies the specific data points the support team needs: order ID, current status, carrier name, tracking number, estimated delivery date, and any delay flags. The assistant is configured to retrieve these data points from the order management system and shipping carrier APIs, then draft a response in English. The integration with Google Workspace connects to Gmail for inbound customer emails and Google Calendar for scheduling follow-ups if a human agent needs to step in. The assistant drafts the response, and a human agent approves it before it is sent. This human-in-the-loop design ensures that any message involving refunds, compensation, or contract changes remains under human control, which is a PCI DSS requirement for payment-related communications. The pilot measures three metrics: first-response time, document turnaround time, and error rate. The baseline is established during the first three days of the pilot, before the assistant is fully active, so the before/after comparison is clean and measurable.

    Delivery Model: Dedicated AI Team for Full-Cycle Deployment

    The dedicated AI team handles the full lifecycle of the pilot. Week one covers the process audit, model deployment on the client’s on-premise hardware, and integration with the order management system and shipping carrier APIs. The team configures the retrieval layer, fine-tunes the model on the client’s historical support tickets, and sets up the Google Workspace integration. Week two is the active pilot period, during which the assistant handles live customer queries under human supervision. The team monitors performance daily, adjusting prompts and retrieval logic as needed. The team also documents the before/after metrics, including first-response time, document turnaround time, and error rate, so the client has a clear measurement of the pilot’s impact. The team operates as an extension of the client’s internal staff, attending daily standups and providing a weekly summary of performance and issues. The client does not need to hire ML engineers or manage infrastructure; the dedicated team handles all technical aspects of the deployment and operation.

    Measured Outcomes: Cycle Time and Error Rate Reduction

    The pilot targets a 60-80% reduction in manual ticket handling for order status queries. First-response time drops from 6 hours to under 2 minutes, because the assistant answers instantly from live data. Document turnaround for shipping confirmations and return authorizations drops from 1-2 business days to same-day processing. The error rate, measured as the percentage of responses that require human correction, is expected to be under 5% after the first week of tuning. The pilot establishes a clear baseline during the first three days, so the before/after comparison is measurable and defensible. If the metrics show a clear improvement, the next phase expands to additional workflows such as returns processing, product recommendations, or bilingual support for German-language queries. The dedicated AI team continues to monitor performance and adjust prompts as the client’s business processes evolve, ensuring that the assistant remains accurate and relevant as the order management system and shipping carrier APIs change.

  • Ticket Triage Automation for UK E-commerce: A 3-Month Fixed-Scope Pilot

    The Problem: Senior Staff Buried in Routine Ticket Triage

    You run a 51-200 person e-commerce operation in the UK. Your support team handles 800 to 1,500 tickets per day across order status, delivery issues, returns, and product questions. Senior staff spend 40-60% of their time on routine triage: reading the ticket, classifying it, routing it to the right queue, and drafting a first response. This work is repetitive, error-prone, and it pulls your most experienced people away from the complex cases that actually need their judgment. The goal is not to replace your support team; it is to free senior staff from routine work so they can focus on escalations, customer retention, and process improvement. The constraint is GDPR: ticket data contains customer names, order numbers, and delivery addresses, so any automation must comply with UK GDPR and the Data Protection Act 2018. The delivery model is a fixed-scope pilot: one ticket category, one helpdesk, one ERP integration, 3 months, measured before/after baselines on cycle time and error rate.

    Prerequisites: What You Need Before Step 1

    Before you write a single line of integration code, you need five things in place. First, a documented list of your top 20 ticket categories with their current routing rules, SLA targets, and escalation paths. This list is your ground truth; without it, the model has no reference for what ‘correct’ routing looks like. Second, API access to your helpdesk (Zendesk, Freshdesk, or similar) and your SAP or Microsoft Dynamics ERP instance. You need read access to order data and write access to ticket status fields. Third, a named GDPR Data Protection Officer or privacy lead who can sign off on the Data Protection Impact Assessment (DPIA). Fourth, a fixed-scope pilot agreement that defines success metrics (cycle time reduction, error rate, cost per ticket), data handling boundaries, and a 3-month timeline. Fifth, a human-in-the-loop approval workflow in your helpdesk UI where agents can accept, edit, or reject the model’s routing suggestion. If any of these are missing, the pilot will stall in week 2 or 3, and you will not have the measured baselines needed to justify scaling.

    Step 1: Audit the Ticket Flow and Define the Baseline

    Run a 2-week process audit on your top 3 ticket categories. Export 500 historical tickets from your helpdesk, tag each one with its final routing destination, cycle time, and error rate (did it go to the wrong queue, get escalated unnecessarily, or take longer than the SLA?). This gives you a baseline: for example, ‘order status’ tickets average 14 minutes from receipt to first response, with a 7% error rate. The audit also reveals which categories are worth automating. If a category has a 90%+ routing accuracy already, the ROI on automation is low. If it has a 40% error rate and a 22-minute cycle time, it is a strong candidate. The output of this step is a one-page brief per category: current metrics, routing rules, and the target metrics for the pilot. This brief becomes the acceptance criteria for the fixed-scope pilot agreement.

    Step 2: Complete the GDPR DPIA and Data Processing Agreement

    Complete a Data Protection Impact Assessment (DPIA) before any ticket data flows through the OpenAI API. The DPIA must document the lawful basis for processing (typically legitimate interest under GDPR Article 6(1)(f)), the categories of personal data involved (names, order numbers, delivery addresses), the retention policy (delete or anonymise ticket payloads after the routing decision is logged), and the security measures (encryption in transit via TLS 1.3, access controls on the API keys). You must also ensure the OpenAI API is covered by a Data Processing Agreement (DPA) with UK Standard Contractual Clauses. If tickets contain health data (e.g., a customer reporting a product caused an injury), Article 9 applies and you need explicit consent or another specific exception. The DPIA is not a one-time document; it must be updated if you change the model, the data flow, or the retention policy. Your DPO signs off on the DPIA before the pilot goes live.

    Step 3: Build the Model-Agnostic Triage Layer

    Build the triage layer as a model-agnostic abstraction. The integration layer calls your helpdesk’s REST API to fetch new tickets, and your SAP or Dynamics ERP’s OData or SOAP endpoints to enrich the ticket with order data (order status, delivery ETA, return eligibility). The enriched ticket payload is sent to the model endpoint, which returns a classification (category, priority, routing destination) and a suggested first-response template. For the pilot, use OpenAI’s GPT-4o API because it requires no GPU infrastructure and provides high-accuracy classification. The model-agnostic design means the routing logic is decoupled from the model: if GDPR or client contracts later demand on-prem inference, you can swap the model endpoint to a locally hosted open-weight model (e.g., Llama 3 70B) without changing the integration layer. The output is written back to the helpdesk via the API, with the model’s confidence score logged for audit.

    Step 4: Integrate with Helpdesk and ERP via API

    Wire the triage layer into your helpdesk and ERP. The helpdesk integration uses the REST API to create a new ticket, update its status, and log the model’s routing decision. The ERP integration uses OData (for Dynamics) or the SAP Business Technology Platform API to fetch order data and update the ticket with order-specific context. The human-in-the-loop approval workflow is critical: the model’s output appears in the helpdesk UI as a suggestion, and a human agent must accept, edit, or reject it before the ticket is routed. For tickets touching money (refunds, chargebacks), health data, or contract terms, the human approval is mandatory and the model’s output is treated as a suggestion only. The approval log feeds back into the model’s prompt engineering in the next sprint, so the system improves over time. This is not a limitation; it is the compliance mechanism that keeps the system within GDPR and internal audit boundaries.

    Step 5: Run the 8-Week Pilot in Three Phases

    Run the pilot in three phases. Phase 1 (weeks 1-2): shadow mode. The model classifies and routes, but a human approves every action. You measure the model’s accuracy against the human-approved outcomes. Phase 2 (weeks 3-6): semi-automated mode. The model handles low-risk categories (e.g., ‘where is my order’, ‘change delivery address’) and escalates the rest to a human. You measure cycle time and error rate for the automated categories. Phase 3 (weeks 7-8): full automation for approved categories with a 5% random sample still routed to a human for quality checks. The pilot ends with a measured before/after report: cycle time reduction (e.g., from 14 minutes to 3 minutes), error rate (e.g., from 7% to 2%), and cost per ticket (e.g., from £4.20 to £2.10). This report becomes the business case for scaling to other departments and ticket categories.