Tag: Cut First-Response Time

  • RAG Assistant for Order Status: 2-Week Pilot in Austrian E-commerce

    The Problem: Manual Order Status Queries in a 25-Person E-commerce Team

    A 25-person e-commerce operation in Vienna handles 400-600 customer inquiries daily, most of them asking where their order is. The support team spends 3-4 hours per agent per day on these repetitive queries, pulling up order management screens, checking carrier tracking numbers, and drafting responses. First-response time averages 6 hours, and document turnaround for shipping confirmations takes 1-2 business days. The business function is customer support, but the bottleneck is manual data retrieval and response drafting, not the actual customer interaction. The need is clear: cut first-response time to under 2 minutes and reduce document turnaround to same-day processing, without adding headcount or replacing existing systems. The solution must work within PCI DSS constraints because the support team occasionally handles refund requests that touch cardholder data, and it must integrate with Google Workspace, which the team already uses for email and calendar management. The pilot scope is one specific workflow: order and shipment status updates, chosen because it is high-volume, rule-based, and has clear before/after metrics to measure success.

    Architecture: Open-Weight Models On-Premise for PCI DSS Compliance

    The architecture uses open-weight models running on the client’s own hardware, not cloud APIs. This is a deliberate choice driven by PCI DSS compliance: cardholder data and transaction details must not leave the client’s controlled infrastructure. The model is a 7B-parameter open-weight variant, fine-tuned on the client’s historical support tickets and order management documentation. It runs on a single GPU server in the client’s data center, with all inference happening locally. The retrieval layer connects to the client’s order management system and shipping carrier APIs via standard REST endpoints, pulling real-time order status, tracking numbers, and delivery windows for each query. The assistant does not store transaction data; it retrieves it on demand, which means the model never has persistent access to sensitive information. This architecture satisfies PCI DSS requirement 3.4, which mandates that cardholder data be rendered unreadable at rest, and requirement 4, which requires encryption of data in transit. The model-agnostic design means that if the client later wants to use a different model for a different workflow, the retrieval layer and integration code remain unchanged.

    Pilot Scope: Two-Week Deployment on Order Status Queries

    The pilot runs for two weeks, starting with a process audit that maps the current workflow for order status queries. The audit identifies the specific data points the support team needs: order ID, current status, carrier name, tracking number, estimated delivery date, and any delay flags. The assistant is configured to retrieve these data points from the order management system and shipping carrier APIs, then draft a response in English. The integration with Google Workspace connects to Gmail for inbound customer emails and Google Calendar for scheduling follow-ups if a human agent needs to step in. The assistant drafts the response, and a human agent approves it before it is sent. This human-in-the-loop design ensures that any message involving refunds, compensation, or contract changes remains under human control, which is a PCI DSS requirement for payment-related communications. The pilot measures three metrics: first-response time, document turnaround time, and error rate. The baseline is established during the first three days of the pilot, before the assistant is fully active, so the before/after comparison is clean and measurable.

    Delivery Model: Dedicated AI Team for Full-Cycle Deployment

    The dedicated AI team handles the full lifecycle of the pilot. Week one covers the process audit, model deployment on the client’s on-premise hardware, and integration with the order management system and shipping carrier APIs. The team configures the retrieval layer, fine-tunes the model on the client’s historical support tickets, and sets up the Google Workspace integration. Week two is the active pilot period, during which the assistant handles live customer queries under human supervision. The team monitors performance daily, adjusting prompts and retrieval logic as needed. The team also documents the before/after metrics, including first-response time, document turnaround time, and error rate, so the client has a clear measurement of the pilot’s impact. The team operates as an extension of the client’s internal staff, attending daily standups and providing a weekly summary of performance and issues. The client does not need to hire ML engineers or manage infrastructure; the dedicated team handles all technical aspects of the deployment and operation.

    Measured Outcomes: Cycle Time and Error Rate Reduction

    The pilot targets a 60-80% reduction in manual ticket handling for order status queries. First-response time drops from 6 hours to under 2 minutes, because the assistant answers instantly from live data. Document turnaround for shipping confirmations and return authorizations drops from 1-2 business days to same-day processing. The error rate, measured as the percentage of responses that require human correction, is expected to be under 5% after the first week of tuning. The pilot establishes a clear baseline during the first three days, so the before/after comparison is measurable and defensible. If the metrics show a clear improvement, the next phase expands to additional workflows such as returns processing, product recommendations, or bilingual support for German-language queries. The dedicated AI team continues to monitor performance and adjust prompts as the client’s business processes evolve, ensuring that the assistant remains accurate and relevant as the order management system and shipping carrier APIs change.

  • UK Medtech Cuts Invoice First-Response Time to 6 Hours with On-Premise AI

    Background: A UK Medtech Distributor at 1,200 Headcount

    This case study is a composite drawn from patterns Forfis has observed across multiple engagements. We do not name real clients. The company described here is a UK-based medtech distributor with roughly 1,200 employees, operating in the 501-2000 band. It handles procurement, supply-chain coordination, and customer-facing service for hospital and clinic clients across the UK and Ireland. The existing stack includes a mid-market ERP, a CRM for customer records, and Microsoft Teams as the primary internal messaging channel. The finance and operations teams were running on a mix of spreadsheets, email threads, and a legacy invoice portal that had not been updated since 2019. The company had no dedicated AI team and had not previously deployed any machine-learning system in production.

    Challenge: 48-Hour Invoice Response, Zero New Hires, GDPR in the Loop

    The trigger was a 40 percent increase in supplier invoice volume over eighteen months, driven by a new product line and expanded distribution contracts. The finance team of eleven was processing invoices manually: extracting line items, matching them against purchase orders, flagging discrepancies, and posting to the ERP. Average first-response time to a supplier query about a disputed invoice was 48 hours. The operations director had a hard constraint: no new headcount in the current fiscal year, and GDPR compliance was non-negotiable because invoice metadata occasionally contained patient-identifiable information from hospital procurement orders. The deadline was six months to show a measurable reduction in cycle time before the next board review. The team needed to cut first-response time without adding a single FTE and without sending regulated data to a third-party cloud API.

    Approach: On-Premise Open-Weight Models, Predictive Scoring, and a Fixed-Scope Pilot

    Forfis ran a two-week process audit across the finance and operations workflows. The audit identified invoice processing as the highest-impact target: high volume, repetitive extraction, and a clear before/after metric. The pilot scope was fixed: one invoice category (supplier purchase orders with line-item extraction), one integration point (the existing ERP API), and one notification channel (Microsoft Teams). The architecture used open-weight models on the client’s own hardware, so no regulated data left the building. A retrieval-augmented layer pulled context from the client’s own procurement documentation and CRM records to improve extraction accuracy. Predictive scoring assigned a confidence value to each extracted field; items above 95 percent auto-posted, items below routed to a human reviewer in Teams. The dedicated AI team of four engineers and one product designer worked on-site for the first four weeks, then shifted to remote with weekly syncs. The pilot ran for eight weeks with a measured baseline captured in week one.

    Outcome: 48 Hours to Under 6, Error Rate Below 2 Percent

    The pilot cleared its threshold. Average first-response time for supplier invoice queries dropped from 48 hours to under 6 hours. Extraction error rate on line items fell from 11 percent to under 2 percent. The finance team’s manual review volume dropped by roughly 60 percent, because the predictive scoring layer auto-approved the high-confidence items. The remaining 40 percent of invoices still required human eyes, but the reviewers now worked from a pre-drafted, context-enriched queue in Teams rather than a blank spreadsheet. The ERP integration held: no data left the client’s infrastructure, and the GDPR data-processing record was updated to reflect the on-premise model deployment. The operations director reported that the team absorbed the 40 percent invoice volume increase without a single new hire. The six-month timeline was met, and the board review proceeded on the strength of the measured baseline.

    Lessons for Similar Teams

    • Baseline first, always. The pilot did not start until the team had a measured before/after baseline on cycle time and error rate. Without that number, the board review would have been a conversation about impressions rather than data. Every similar team should capture the baseline in week one, not after the pilot ends.
    • Model-agnostic architecture pays off. The client started with open-weight models on-premise for GDPR reasons. If a future use case requires a frontier API for a non-regulated workflow, the integration layer does not need to be rebuilt. Teams that hard-code a single vendor API into their architecture will face this problem.
    • Predictive scoring is the human-in-the-loop mechanism. The confidence threshold is not a suggestion; it is the architectural gate. Items above 95 percent auto-approve, items below route to a human. This is what makes GDPR Article 22 compliance operational rather than theoretical.
    • Integration through existing APIs, not replacement. The ERP, CRM, and Teams stack stayed intact. The AI layer sat on top. For a 1,200-person operation, a rip-and-replace project would have taken two years and a budget the company did not have.
    • Dedicated team beats rotating contractors. The four engineers and one product designer stayed on the engagement from audit through rollout. Consistency in the team meant the client’s internal stakeholders had a single point of contact and a shared context that did not reset every sprint.
  • Cutting First-Response Time in UK Logistics: A 4-Week AI Ticket Triage Pilot

    The Problem: Slow First-Response Time in UK Logistics Support

    You run a 500-to-2,000-person logistics or supply chain operation in the UK. Your customer support team handles 800 to 3,000 tickets per week across email, web forms, and a helpdesk portal. First-response time sits at 4 to 12 hours, and 30 to 50 percent of tickets are misrouted to the wrong team, forcing manual reassignment. You have run isolated AI pilots before — perhaps a document extraction proof-of-concept or a chatbot experiment — but none have moved into production. Your ISO 27001 certification requires that any new system touching customer data passes a formal risk assessment, and your operations team needs a measured before/after baseline on cycle time and error rate before approving rollout. The goal is not to replace your support staff but to cut first-response time by 30 to 50 percent within four weeks, using predictive scoring to route tickets to the correct team before a human ever opens them.

    Prerequisites Before You Start

    Before you write a single line of integration code, confirm these items are in place:

    • Process map: A documented flow of how tickets currently move from intake to resolution, including which teams handle which categories (delivery delays, billing disputes, customs queries, returns).
    • API credentials: Read/write access to your helpdesk (Zendesk, Freshdesk, Jira Service Management) and CRM via their REST APIs. You will need webhook endpoints for real-time ticket events.
    • ISO 27001 owner: A named compliance lead who can sign off on the risk assessment for using OpenAI API with customer data. This person must be involved from Day 1, not after the pilot is built.
    • Pilot budget: £1,500 to £4,000 for OpenAI API costs over four weeks, plus £8,000 to £15,000 for fixed-scope integration work. Confirm this with finance before Week 1 starts.
    • Operations lead: One person with 5 to 10 hours per week to review model outputs, approve routing rules, and flag misrouted tickets during the pilot.
    • Data samples: 200 to 500 historical tickets with metadata (sender, category, resolution time, team assigned) to train and validate the scoring model.

    Step-by-Step: Build the Pilot in Four Weeks

    Step 1: Run the process audit and capture baselines. Map every ticket category, the team that handles it, and the average time from intake to first response. Export 200 to 500 historical tickets from your helpdesk with fields: ticket_id, sender_email, subject, body, assigned_team, first_response_time_hours, resolution_time_hours, category. Store this in a CSV or database table. This is your before-state. Without it, you cannot prove the pilot worked.

    Step 2: Define routing categories and scoring thresholds. List 5 to 8 ticket categories your support team actually uses (e.g., delivery_delay, billing_dispute, customs_query, return_request, account_issue). For each, define what a correct routing looks like. Set a confidence threshold: tickets scoring 0.85 or above are auto-routed; below 0.85 go to a human queue. Document this in a one-page routing spec that your ISO 27001 owner signs off.

    Step 3: Build the OpenAI API integration via REST and webhooks. Create a webhook listener in your helpdesk that fires on ticket.created. The listener sends the ticket body and metadata to a lightweight service (Node.js or Python) that calls the OpenAI API using the gpt-4o-mini model. The prompt instructs the model to return a JSON object: {"category": "delivery_delay", "confidence": 0.92, "suggested_team": "dispatch"}. Log every API call with timestamp, ticket ID, and response in your SIEM to satisfy ISO 27001 Annex A.12.3.1.

    Step-by-Step: Run the Pilot and Hand Over

    Step 4: Implement human-in-the-loop approval. Any ticket with a confidence score below 0.85, or any ticket mentioning financial amounts, health data, or contract terms, is flagged for human review. Build a simple approval screen in your helpdesk or a lightweight web app where the operations lead sees the AI’s suggested routing, can accept or override it, and logs the reason for any override. This is not optional under ISO 27001 — you must demonstrate that a human controls decisions touching money or regulated data.

    Step 5: Run the pilot on live tickets for two weeks. Enable the webhook on 100 to 200 live tickets per day. The AI scores and routes; the operations lead reviews every ticket for the first three days, then samples 20 percent after that. Track daily: first-response time, routing accuracy (correct team vs. AI suggestion), override rate, and API cost. If the override rate exceeds 15 percent in any 7-day window, pause the pilot and recalibrate the prompt or scoring thresholds.

    Step 6: Measure before/after and document findings. In Week 4, compare the pilot metrics against your Week 1 baselines. You should see first-response time drop by 30 to 50 percent and routing accuracy at 85 percent or above. Write a two-page report: what worked, what failed, API costs, and a recommendation for rollout. This report is your input to the ISO 27001 management review and your business case for scaling to additional teams or channels.

    Step 7: Hand over to managed operations. If the pilot meets targets, transition to a managed operations model. Forfis continues to monitor model performance, tune routing thresholds monthly, update prompts as new ticket patterns emerge, and handle API cost management. You retain ownership of the data and the integration; Forfis operates the AI layer under a service-level agreement with defined accuracy and latency targets.

    Common Pitfalls and How to Detect Them

    • Overfitting on historical patterns: The model learns routing rules from last year’s ticket mix, but your operations have changed (new routes, new clients, new service levels). Detect this by tracking the override rate weekly. If it climbs above 15 percent, the model is misrouting. Recalibrate by retraining on the last 30 days of tickets, not the full historical set.

    • Skipping the human-in-the-loop step for high-value tickets: You auto-route a billing dispute because the confidence score is 0.87, but the ticket involves a £50,000 claim. This is an ISO 27001 breach. Detect this by auditing the approval log monthly. Any ticket with a financial amount above your defined threshold (e.g., £1,000) must have a human approval record.

    • Ignoring API cost creep: GPT-4o-mini costs roughly £0.15 per 1,000 input tokens and £0.60 per 1,000 output tokens. A 500-word ticket with a 200-word response costs about £0.05. At 2,000 tickets per week, that is £500 per week. If you do not set a monthly API budget cap in your OpenAI dashboard, costs can double if ticket volume spikes during peak season. Detect this by reviewing API spend weekly against your pilot budget.

    • Not logging API calls for ISO 27001 audit: If you do not log every OpenAI API call with timestamp, ticket ID, and response, you cannot demonstrate compliance during an ISO 27001 surveillance audit. Detect this by running a monthly audit of your SIEM logs. If any ticket ID is missing from the log, the integration is not compliant.

    What Comes After the Pilot

    The pilot is not the end state. Once you have a measured before/after baseline and a signed-off ISO 27001 risk assessment, the next logical step is to extend the triage layer to additional channels — voice, chat, or email — and to add document extraction for attached invoices, customs forms, or proof-of-delivery images. The same predictive scoring architecture applies: the model classifies the document type, extracts key fields, and routes the data to your ERP or accounting system. The human-in-the-loop control remains for anything touching money or regulated data. Your four-week pilot gives you the data, the compliance sign-off, and the operational muscle to justify that next phase to your board or investors. The integration is already built; the next step is scaling it.

  • Cutting First-Response Time from 38 Hours to 4 in a German Medtech Distributor

    Background: A Mid-Size Medtech Distributor in Southern Germany

    This case study is a composite. It draws on patterns Forfis has observed across multiple engagements in German healthcare and medtech distribution. No named customer is represented; the company, metrics, and timeline are representative of the work we deliver, not a single identifiable client.

    The company in question is a mid-size medtech distributor in southern Germany, roughly 340 employees, operating across two regional warehouses and a central back office in Stuttgart. It handles order intake, shipment coordination, and after-sales support for orthopedic and diagnostic equipment. The ERP is SAP S/4HANA, the helpdesk is a legacy on-premises ticketing system, and the CRM is Microsoft Dynamics 365. The company had been running on a paper-and-email hybrid for inbound purchase orders and shipment confirmations for over a decade. No prior AI or automation project had been attempted; the operations team had flagged the bottleneck in internal reviews for three consecutive quarters without a funded solution.

    Challenge: A 24-Hour SLA the Manual Process Could Not Meet

    The trigger was a contractual deadline. A major hospital group, representing roughly 18 percent of the company’s annual revenue, issued a service-level agreement requiring order-status acknowledgments within 24 hours and shipment confirmations within 4 hours of dispatch. The existing process could not meet either threshold. Inbound purchase orders arrived as scanned PDFs, emailed attachments, and occasionally physical mail. A team of four operators manually transcribed each order into SAP, cross-referenced it against the shipment plan, and drafted a status email to the customer. The median first-response time was 38 hours. The 95th percentile was 72 hours. The error rate on transcribed fields was 6.2 percent, and each correction required a second pass through the approval chain.

    The operational pressure was compounded by GDPR. The documents contained patient identifiers, billing addresses, and in some cases clinical context. The company’s data protection officer had flagged the manual process as a compliance risk: paper documents were stored in unsecured filing cabinets, and email attachments were not consistently encrypted. The deadline was not optional. The hospital group had indicated that non-compliance would trigger a contract review in the following quarter.

    Approach: A Six-Week Integration Sprint on SAP and Claude

    Forfis ran a six-week integration sprint. The first week was a process audit: mapping every document type, every handoff, every approval gate, and every data field that touched the ERP. The audit identified 14 distinct document formats across purchase orders, packing lists, customs declarations, and shipment confirmations. The team selected the three highest-volume formats for the pilot, covering roughly 70 percent of inbound documents.

    The extraction pipeline used the Anthropic Claude API for document parsing and field classification. The model was prompted with structured output schemas matching the SAP data model. The orchestration layer, built on a workflow engine, routed each extracted record through a confidence check. Records above a 92 percent confidence threshold and containing no patient identifiers or payment amounts were auto-approved. Everything else went to a human approver in a queue built into the existing helpdesk. The SAP integration used the OData API to write order and shipment records directly into S/4HANA, bypassing the manual entry step entirely. The first-response template engine pulled the enriched record from SAP and generated a status email within 90 seconds of approval.

    Outcome: 38 Hours to 4 Hours, 6.2 Percent to 0.9 Percent

    The pilot went live in week seven on a subset of order types from two regional warehouses. The full rollout followed in weeks eight and nine, extending to all document types and both warehouses. The two-month stabilization phase that followed focused on reducing the human-review rate and tuning extraction thresholds per document type.

    The measured outcomes, tracked against the pre-pilot baseline, were as follows:

    • Median first-response time fell from 38 hours to 4 hours. The 95th percentile dropped from 72 hours to 11 hours.
    • Error rate on extracted fields fell from 6.2 percent to 0.9 percent.
    • Cycle time per document, from receipt to ERP entry, dropped from 4.5 hours to 22 minutes.
    • Human-review rate settled at 15 to 20 percent of records in steady state, down from the initial 25 percent.
    • Customer satisfaction for order-status inquiries rose by 11 points on a 100-point scale over the first quarter after go-live.

    Two full-time operators were redirected from manual data entry to exception handling and quality review. The GDPR compliance work, including the DPIA under Article 35 and the pseudonymization pipeline, was completed before go-live and required no rework during the stabilization phase.

    Lessons for Similar Teams in Healthcare and Medtech

    Five lessons from this engagement generalize to similar teams in healthcare and medtech distribution:

    • Start with the SLA, not the technology. The hospital group’s 24-hour acknowledgment requirement defined the success criterion. The technology choice followed from the constraint, not the other way around. Teams that start with a model demo and work backward to a business need tend to over-build and under-deliver.

    • The process audit is not optional. The 14 document formats, the unsecured filing cabinets, the inconsistent email encryption — none of this was visible from a technology specification. The audit took one week and saved an estimated three weeks of rework later in the sprint.

    • Human-in-the-loop is a design decision, not a fallback. The confidence threshold and the data-sensitivity routing were defined in week two, before any code was written. Teams that treat the human gate as an afterthought end up with either over-automation (errors in production) or under-automation (the human reviews everything, and the cycle time does not improve).

    • Model-agnostic architecture protects the client’s future. The client’s data protection officer asked, in week four, whether the pipeline could run on an open-weight model if the hospital group’s contract was renegotiated. Because the orchestration layer was decoupled from the model API, the answer was yes, and the rework estimate was under two weeks. A hard-coded dependency on a single vendor API would have made that conversation much harder.

    • The baseline is the deliverable. The before/after measurement on cycle time and error rate was agreed in the audit phase and tracked from day one of the pilot. Without that baseline, the 38-to-4-hour improvement would have been anecdotal. With it, the client could present the numbers to the hospital group’s procurement team with confidence.

  • Cutting First-Response Time on Order-Status Tickets with LangGraph and RAG

    The Problem: Serial Ticket Handling in High-Volume E-commerce Support

    A 2,000+ employee e-commerce company in the USA handles roughly 50,000 support tickets per month. A significant share of those are order and shipment status inquiries: “Where is my package?” “Why is my order delayed?” “I haven’t received my confirmation email.” Each one lands in a shared Gmail inbox, gets picked up by an agent, who logs into the order management system, checks the shipment tracker, drafts a reply, and sends it. Average first-response time sits at 4-6 hours during peak season, and the cost per ticket is driven almost entirely by agent labor.

    The problem is not that agents are slow. It is that the workflow is serial: a human must read the ticket, decide what data to pull, pull it from two or three systems, compose a response, and send it. The AI opportunity is not to replace the agent but to collapse the serial steps into a parallel pipeline where the machine does the retrieval and drafting, and the human does the approval. Forfis approaches this as a workflow orchestration problem, not a chatbot problem. The goal is to cut first-response time from hours to minutes while keeping a human in the loop for anything that touches money or a customer commitment.

    The Mechanism: LangGraph Orchestration with a RAG Retrieval Layer

    The architecture rests on three layers. The orchestration layer uses LangGraph to define a stateful graph where each node is a discrete step: classify the ticket, retrieve order data, draft a response, check the approval gate, and send. Edges between nodes encode the control flow, including branches for escalation to a human agent when confidence is below threshold. LangChain sits underneath, providing the abstractions for LLM calls, prompt management, and document retrieval.

    The retrieval layer is a RAG pipeline. The company’s order management system, shipment tracking data, and policy documents are chunked at the record level and embedded into a vector store. When a ticket arrives, the system retrieves the relevant order record and passes it as context to the LLM. The integration layer connects to Google Workspace via the Gmail API and Google Chat API using OAuth 2.0 with least-privilege scopes. The AI does not replace the mailbox; it drafts responses that a human agent reviews and sends through the existing interface.

    The model choice is deliberately model-agnostic. Classification and retrieval run on an open-weight model on the client’s hardware where data residency matters. Final response drafting uses a frontier API (OpenAI or Anthropic) for quality. LangGraph abstracts this, so swapping models does not require re-architecting the graph.

    Trade-offs: Latency, Data Residency, and Automation Depth

    The first trade-off is latency versus accuracy. A frontier API produces better-drafted responses but adds 1-3 seconds of network latency per call. For a first-response-time target of under 10 minutes, this is acceptable. For a real-time voice channel, it would not be. The second trade-off is data residency versus model quality. Running the RAG pipeline on an open-weight model on-premises keeps customer order data inside the building, satisfying ISO 27001 data classification controls, but the model’s drafting quality is lower than a frontier API. The hybrid approach — on-premises retrieval, cloud drafting — splits the difference.

    The third trade-off is automation depth versus risk. Auto-approving every AI-drafted response would cut first-response time to under 2 minutes, but it violates the human-in-the-loop requirement for anything touching a refund or a contract. Forfis sets the approval gate at the record level: routine order-status queries auto-approve above a confidence threshold, but any response that mentions a refund, a delay compensation, or a policy exception routes to a human. This keeps the 90% of tickets that are simple status checks fast while protecting the 10% that carry financial or legal risk.

    The fourth trade-off is integration scope versus timeline. A four-week sprint cannot rebuild the CRM or the order management system. The integration is read-only on the data sources and write-only on the Gmail outbox. This constraint is a feature: it keeps the pilot reversible and the blast radius small.

    Recommendation: Start with a Fixed-Scope Pilot on Order-Status Tickets

    For a 2,000+ employee e-commerce company in the USA targeting ISO 27001 compliance, the recommendation is to start with a fixed-scope pilot on order and shipment status tickets only. Do not attempt to automate refund processing, returns, or policy exceptions in the first sprint. The pilot should measure three baselines before the AI goes live: average first-response time, average handling time, and error rate (wrong order number cited, incorrect shipment status, policy misstatement). After four weeks, compare the post-pilot numbers against the baseline.

    The integration sprint should follow this sequence: Week one is the process audit and baseline measurement. Weeks two and three build the LangGraph graph, wire the RAG pipeline to the order and shipment data, and connect the Google Workspace API. Week four is the pilot with the human-in-the-loop gate active. The pilot ships with a documented before/after report on cycle time and error rate.

    Two specific recommendations. First, chunk the RAG index at the record level, not the paragraph level. Order data is structured; the LLM needs the full order record to answer accurately. Second, log every AI-drafted response, every retrieval, and every approval decision. ISO 27001 requires documented evidence of information security controls, and the audit log is that evidence. The log should capture the ticket ID, the retrieved records, the model used, the confidence score, and the approver’s identity. This log is also the foundation for the managed operation phase after the pilot.

  • Cutting First-Response Time in a 51-200-Person B2B SaaS: A 2-Week pgvector Pilot

    The First-Response Bottleneck in a 51-200-Person B2B SaaS Team

    A 51-200-person B2B SaaS company in Germany runs 40-120 inbound leads per week across forms, chat, and email. The sales and marketing teams handle triage manually: a person reads each submission, checks the CRM for duplicates, looks up the prospect’s company in a spreadsheet, and drafts a first response. Cycle time averages 12-48 hours. Error rate on lead classification sits at 15-25% because the team works from memory and inconsistent notes. The marketing team maintains product docs in Notion or Confluence, but sales reps rarely reference them when writing replies, so answers drift from the official positioning.

    The pain is not a lack of effort. It is a structural mismatch: the team has 6-10 people covering sales, marketing, and support, and the volume of inbound leads grows 15-20% quarter-over-quarter. Hiring two more SDRs costs EUR 120,000-160,000 per year in salary and benefits, and the new hires need 8-12 weeks to reach full productivity. The existing team is already at capacity, and the first-response metric is slipping because the queue grows faster than the headcount.

    Why Generic Chatbots and Rule-Based Workflows Fall Short

    Most teams reach for a generic chatbot or a rule-based CRM workflow. The chatbot answers from a fixed FAQ, so it cannot reference the specific product doc a prospect just read or the integration they asked about. The rule-based workflow tags leads by form field, but it does not enrich the record with firmographic data or clean up inconsistent CRM entries. Both approaches reduce manual effort but do not cut first-response time below 4 hours because the human still drafts the reply from scratch.

    A second common approach is to hire a junior SDR to handle triage. This works until the lead volume doubles, and the junior SDR becomes the new bottleneck. The cost scales linearly with volume, and the quality of classification depends on the individual’s familiarity with the ICP, which varies by day. Neither approach addresses the root problem: the team lacks a system that grounds responses in the company’s own documentation and enriches the CRM record automatically.

    The failure mode is not the technology. It is the architecture. A chatbot without retrieval-augmented generation cannot answer questions that require context from your specific docs. A rule-based workflow without data enrichment leaves the CRM record incomplete, so the next step in the sales process starts from a blank slate.

    A pgvector-Grounded Assistant That Qualifies Leads and Enriches CRM Data

    The approach starts with a process audit that maps the lead-qualification workflow end to end: form submission, CRM entry, duplicate check, firmographic lookup, classification, first-response drafting, and human approval. The audit identifies the two highest-leverage steps: drafting the first response and enriching the CRM record. The pilot targets those two steps on one workflow, typically the primary inbound form, and runs for 2 weeks.

    The architecture uses pgvector embeddings search to ground the assistant in the company’s own documentation. The system ingests Notion or Confluence pages via API, chunks them into 256-512 token segments, embeds them, and stores the vectors in pgvector. When a lead asks a question, the system embeds the query, retrieves the top 5-10 most relevant chunks, and feeds them to the LLM as context. The LLM composes a response that cites the source doc, so the answer reflects the current positioning rather than the model’s training data.

    The model layer is deliberately model-agnostic. For high-quality drafting and classification, the system uses OpenAI or Anthropic APIs hosted in EU data centers to satisfy GDPR data-residency requirements. For regulated data that cannot leave the building, the system runs an open-weight model on the client’s own hardware. The integration layer plugs into the existing CRM, helpdesk, and messaging tools through their APIs, so no system is replaced. The delivery model is managed AI operations: the team monitors model performance, re-indexes embeddings when docs change, tunes prompts, and handles GDPR compliance checks on an ongoing basis.

    How to Start: A 2-Week Pilot on One Workflow

    Week 1: Run the process audit. Map the lead-qualification workflow, measure the baseline cycle time and error rate over 2 weeks of historical data, and identify the two highest-leverage steps. The audit takes 3-5 days and produces a one-page summary with specific numbers.

    Week 2: Build the pilot. Ingest the Notion or Confluence workspace, chunk and embed the docs, and store the vectors in pgvector. Connect the CRM via API so the assistant can read and write lead records. Configure the LLM to draft first responses grounded in the retrieved chunks. Set up the human-in-the-loop approval step: the assistant drafts, a person reviews and approves before the reply goes out.

    Week 3-4: Run the pilot. The assistant handles all inbound leads on the primary form. Measure cycle time, error rate, and first-response time against the baseline. At the end of 2 weeks, produce a before/after report with specific metrics. If the numbers justify it, extend the pilot to additional workflows and departments under a managed operations contract.

    Pitfalls to Avoid in the First 30 Days

    The most common pitfall is skipping the baseline measurement. Without a 2-week pre-pilot baseline on cycle time and error rate, the team cannot prove the pilot worked. The second pitfall is ingesting the entire Notion or Confluence workspace without chunking. Large documents produce noisy embeddings, and the retrieval step returns irrelevant chunks. Chunking into 256-512 token segments with a 50-token overlap improves retrieval precision by 20-30%.

    The third pitfall is ignoring GDPR from the start. The system must log every data access, support right-to-erasure requests by purging embeddings and raw records from pgvector and the CRM, and process personal data only within EU data centers. The data-processing agreement must cover the AI vendor, the vector store, and the integration layer. If the team adds GDPR compliance after the pilot, the rework takes 2-3 weeks and delays rollout.

    The fourth pitfall is treating the pilot as a one-time project. The managed operations contract is not optional. The embedding index degrades as docs change, the LLM API updates its model versions, and the CRM schema evolves. Without ongoing monitoring and re-indexing, the assistant’s accuracy drops within 6-8 weeks, and the team loses trust in the system.

  • GDPR-Compliant AI Lead Qualification Pilot for Swiss Logistics

    The Problem: Slow Lead Qualification Under GDPR Constraints

    A 2000+ employee logistics firm in Switzerland handles 4,000 inbound leads per month across email, web forms, and Slack. Sales reps spend 18 minutes per lead on manual qualification, and first-response time averages 4.2 hours. GDPR Article 22 restricts automated decision-making with legal or similarly significant effects, so any AI that influences contract terms or pricing must keep a human in the loop. The goal is to cut first-response time to under 30 minutes while staying compliant. The pilot targets one workflow: lead qualification. It uses a conversational agent with pgvector embeddings search over the firm’s CRM records and documentation, integrated into Slack or Microsoft Teams. The architecture is model-agnostic, using open-weight models on local hardware where regulated data cannot leave the building.

    Prerequisites: What You Need Before Starting

    Before step 1, you need the following in place:

    • API access to the CRM (e.g., Salesforce, HubSpot) and Slack or Microsoft Teams, with webhook configuration enabled.
    • Data inventory: a list of all documents, CRM fields, and Slack channels the agent will access, mapped to GDPR Article 30 records.
    • Baseline metrics: current first-response time, cycle time, and error rate for lead qualification, measured over at least 30 days.
    • Legal sign-off: confirmation from the DPO that the pilot complies with GDPR Article 6 (lawful basis) and Article 22 (automated decision-making).
    • Hardware: if using open-weight models, a GPU server with at least 24 GB VRAM on the client’s own network.
    • Team: a product owner, a technical lead, and a compliance officer available for weekly check-ins.

    Step 1: Audit the Lead Qualification Workflow

    Run a process audit on the lead qualification workflow. Map every step from inbound lead to qualified opportunity. Identify where manual work occurs: data entry, document extraction, classification, and response drafting. Measure cycle time and error rate for each step. For a logistics firm, typical bottlenecks include manual CRM data entry (12 minutes per lead) and inconsistent qualification criteria across reps. The audit output is a prioritized list of automatable steps, with the top candidate selected for the pilot. This step takes 1-2 weeks and requires access to the CRM and Slack or Teams logs.

    Step 2: Build the pgvector RAG Pipeline

    Build the pgvector index over the firm’s documentation and CRM records. Export relevant documents (pricing sheets, service descriptions, past lead records) into a PostgreSQL table with a vector column. Use an embedding model such as text-embedding-3-small (OpenAI) or bge-large-en-v1.5 (open-weight) to generate 1,536-dimensional vectors. Create an HNSW index with m=16 and ef_construction=64 for fast similarity search. For 50,000 documents, top-5 retrieval should return in under 15 ms. Store metadata (document ID, source, last updated) alongside each vector for audit trails. This step takes 1-2 weeks and requires a PostgreSQL instance with the pgvector extension installed.

    Step 3: Develop the Conversational Agent with Human-in-the-Loop

    Develop the conversational agent that drafts responses and classifies leads. The agent receives an inbound lead via Slack or Teams webhook, retrieves the top-5 relevant chunks from pgvector, and injects them into the prompt. The LLM generates a draft response and a qualification score (e.g., 1-10) based on the context. The agent posts the draft to a designated Slack channel or Teams channel for human review. A sales rep approves, edits, or rejects the draft. Every interaction is logged with timestamps, the retrieved context, and the final decision. The agent uses OpenAI or Anthropic APIs for quality, or open-weight models on local hardware for regulated data. This step takes 2-3 weeks.

    Step 4: Run the Fixed-Scope Pilot

    Run the fixed-scope pilot on the lead qualification workflow for 4-6 weeks. The agent handles all inbound leads in the designated Slack or Teams channel. A sales rep reviews and approves every draft. Measure first-response time, cycle time, and error rate daily. Compare against the baseline from the audit. For a logistics firm, the target is to cut first-response time from 4.2 hours to under 30 minutes and reduce error rate from 12% to under 5%. Log every interaction for GDPR audit trails. If the agent’s qualification score diverges from the human’s decision by more than 2 points, flag it for review. This step takes 4-6 weeks and requires daily monitoring.

    Step 5: Measure, Refine, and Document the Rollout Plan

    Analyze the pilot results and document the rollout plan. Compare before/after metrics on cycle time, error rate, and first-response time. Identify failure modes: cases where the agent’s draft was rejected, where the qualification score was wrong, or where the retrieved context was irrelevant. Refine the prompt, the pgvector index, or the approval threshold based on the findings. Document the rollout plan for the next phase: multi-channel integration, full CRM sync, and managed operation. The deliverable is a measured baseline, a refined agent, and a clear path to scale. This step takes 1-2 weeks and requires a review meeting with the product owner, technical lead, and compliance officer.

  • Cutting First-Response Time in Swiss Insurance Hiring with a LangGraph Pilot

    The 48-Hour Black Hole in Swiss Insurance Hiring

    A 501-2000 employee insurer in Switzerland receives 300-500 applications per week across 15-20 open roles. Recruiters manually triage each CV, score it against a rubric, and draft a response. The median time-to-first-response is 48-72 hours. Candidates who do not hear back within 48 hours are 3x more likely to accept a competing offer. The recruiter team is flat: no new hires are planned for the next 12 months. The operations team is asked to cut first-response time without adding headcount. The constraint is not technical; it is structural. The current process is a linear, human-bottlenecked pipeline that cannot scale with application volume.

    Why Off-the-Shelf ATS and In-House ML Both Fail

    The first common approach is to buy an off-the-shelf ATS with an AI scoring module. These tools parse CVs and assign a score, but the scoring rubric is opaque and not configurable to the insurer’s specific role requirements. The second approach is to build a custom ML model in-house. This takes 6-12 months, requires a data science team the insurer does not have, and produces a model that is hard to audit under the EU AI Act. The third approach is to outsource to a staffing agency. This reduces recruiter workload but does not cut first-response time; the agency’s own triage process is equally slow. None of these approaches address the root cause: the workflow is not orchestrated. It is a sequence of manual steps with no state management, no branching logic, and no audit trail.

    A LangGraph Workflow with Human-in-the-Loop Approval

    The alternative is a workflow-orchestration approach built on LangChain and LangGraph. LangChain provides the abstraction layer for calling LLMs, vector stores, and tools. LangGraph adds a stateful, cyclic execution model where each node is a function (e.g., ‘parse CV’, ‘score against rubric’, ‘flag for human review’) and edges define control flow. For candidate screening, the workflow is a DAG: the CV is ingested from Google Workspace (Gmail API), parsed into structured data, scored against a predefined rubric, and routed to a human-approval gate if the score is borderline. The AI drafts the response email; the recruiter approves it before it is sent. The architecture is model-agnostic: open-weight models on the client’s own hardware where CVs contain health or financial data, commercial APIs where quality matters. The output is a measured before/after baseline on cycle time and error rate, shipped in a 2-week pilot.

    The 2-Week Pilot: Audit, Build, Measure

    The pilot is scoped to 50-100 real candidates over two weeks. Week 1: the AI process audit maps the current screening steps, identifies the 2-3 highest-volume, lowest-complexity tasks, and selects the LLM. The LangGraph workflow is built with a human-approval gate and a logging mechanism that captures every decision. Week 2: the pilot runs on live applications. The team measures median time-to-first-response, error rate in CV parsing, and recruiter time saved. The output is a go/no-go decision for scaling to all hiring pipelines. The managed AI operations model means the workflow is monitored, tuned, and updated after the pilot; the insurer does not own the maintenance burden. The EU AI Act compliance artifacts (risk management documentation, technical documentation, oversight logs) are produced as part of the pilot, not as a separate project.

    Five Concrete First Steps

    The first step is to define the success metric: median time-to-first-response, not average. The second is to establish the baseline: manually track 50-100 applications for one week before the pilot. The third is to scope the pilot: select the 2-3 highest-volume roles, define the scoring rubric (5-7 criteria), and identify the human-approval gate. The fourth is to choose the LLM: open-weight on-prem if CVs contain regulated data, commercial API otherwise. The fifth is to build the LangGraph workflow with a logging mechanism that captures every decision for the EU AI Act compliance file. The pilot is not a proof of concept; it is a measured, compliance-ready baseline that the insurer can use to justify scaling to all hiring pipelines.

  • How a 120-Person UK Advisory Firm Cut Contract First-Response Time to 38 Minutes

    Background: A 120-Person UK Advisory Firm

    This case study is a composite based on patterns Forfis has observed across multiple engagements in the UK professional services sector. No named client is represented; the figures are drawn from real pilot baselines and post-rollout measurements. The company in this story is a 120-person firm providing legal and financial advisory services to mid-market clients in London and Manchester. It runs on Microsoft 365, a mid-tier CRM, and a document management system that predates the current team. The firm sits in the 51-200 employee band, which means it has the volume to justify automation but not the headcount to run a dedicated AI team.

    The Challenge: 4.2-Hour First Response and a 14-Week Deadline

    The firm’s contract review process was the bottleneck. Clients sent contracts via email; a paralegal or junior associate extracted key clauses, flagged risks, and drafted a response. First-response time averaged 4.2 hours, with a peak of 11 hours during quarter-end. The error rate on clause extraction was 6.1%, meaning roughly one in sixteen contracts required a second pass. GDPR Article 22 required that no automated system make a decision solely on the basis of profiling without human oversight. The firm also faced a deadline: a major client contract was due in 14 weeks, and the existing team could not absorb the volume without hiring two additional paralegals at a cost of approximately GBP 78,000 per year.

    Approach: Audit, Pilot Sprint, and Model-Agnostic Integration

    Forfis began with a two-week process audit. The team mapped every step of the contract review workflow, measured cycle time and error rate on a sample of 200 contracts, and scored each sub-task by volume, error rate, and regulatory exposure. The audit produced a phased roadmap: a fixed-scope pilot on clause extraction and risk flagging, followed by rollout to the financial advisory team. The pilot used the Anthropic Claude API for extraction and classification, with a human-in-the-loop approval gate for anything touching contract terms. The integration sprint ran five weeks: Forfis built the extraction pipeline, connected it to the firm’s CRM and Microsoft Teams, and shipped a Slack channel where flagged clauses appeared as threaded messages with confidence scores. The model-agnostic architecture meant the firm could swap to an open-weight model on its own hardware if data residency requirements tightened.

    Outcome: 38-Minute First Response and a 1.4% Error Rate

    After the five-week pilot, the firm measured the new baseline. First-response time dropped from 4.2 hours to 38 minutes. Extraction error rate fell from 6.1% to 1.4%. The paralegal team redirected its time from manual extraction to higher-value risk analysis. The firm did not hire the two additional paralegals. Rollout to the financial advisory team took three additional weeks, extending the total engagement to three months. The managed operation phase began in week 13, with Forfis monitoring model performance, handling edge cases, and tuning the extraction prompts. The client retained ownership of the integration code and the Teams/Slack configuration, so it could extend the workflow internally without a new engagement.

    Lessons for Similar Teams

    • Baseline before you build. The audit’s 200-contract sample gave the firm a defensible before/after metric. Without it, the pilot’s success would have been anecdotal. Teams that skip the baseline struggle to justify scaling to stakeholders.
    • Human-in-the-loop is not a compromise. The approval gate for contract terms kept the firm compliant with GDPR Article 22 while still cutting manual effort. The gate added 12 seconds per clause but prevented a single high-risk auto-approval that would have required a client call.
    • Model-agnosticism is a risk hedge. The firm’s data residency requirements could have shifted mid-engagement. Because the architecture supported open-weight models on local hardware, Forfis could swap the backend without rewriting the integration layer.
    • Integration over replacement. Plugging into the existing CRM and Teams meant the team did not have to learn a new tool. Adoption was near-complete in the first week because the workflow appeared in the channel they already checked every morning.
    • Fixed-scope pilots reduce scope creep. The five-week sprint had a defined set of document types and a defined approval gate. Adding new document types was a separate decision, not a mid-sprint change request.
  • Voice Agent for Ticket Triage in Fintech: A 4-Week Audit and Pilot

    The Problem: First-Response Time and Cost Per Ticket

    A 501-2000 employee fintech company in the USA handles 12,000 support tickets per month. The average first-response time is 4.2 hours, and the cost per ticket is $18. The company’s support team is stretched thin, and the first-response time is a key driver of customer churn. The company has tried to cut costs by hiring more support agents, but the cost per ticket has not decreased. The company has also tried to use a commercial AI assistant, but the assistant is not PCI DSS compliant and cannot handle card numbers. The company needs a solution that is PCI DSS compliant, can handle card numbers, and can cut the first-response time and the cost per ticket. The solution is a voice agent that is built on an on-premise model and integrated with the company’s CRM, helpdesk, and Notion/Confluence. The voice agent is built by Forfis, a product studio with eight years of delivery experience. The voice agent is built in 4 weeks, and the cost per ticket is cut by 40%.

    The Mechanism: On-Premise Models and Voice Agent Architecture

    The voice agent is built on an on-premise model, which is a Llama 3 70B model. The model is fine-tuned on the company’s ticket data, which includes the ticket category, the ticket priority, and the ticket resolution. The model is stored on the company’s hardware, and the model is updated quarterly. The voice agent uses a speech-to-text model to transcribe the call, a language model to classify the ticket, and a text-to-speech model to generate the response. The speech-to-text model is open-weight and runs on the company’s hardware. The language model is also open-weight and runs on the company’s hardware. The text-to-speech model is a commercial API, because the quality of the voice is important for customer-facing interactions. The integration with the CRM and helpdesk is via their APIs, which are well-documented and stable. The integration with Notion/Confluence is a read-only integration that pulls the company’s documentation into the agent’s context.

    The Trade-Offs: On-Premise vs. Commercial APIs

    The trade-off between on-premise models and commercial APIs is a key decision in the architecture. On-premise models are more expensive to build and maintain, but they are more secure and more compliant. Commercial APIs are cheaper to build and maintain, but they are less secure and less compliant. For a fintech company that is PCI DSS compliant, the on-premise model is the right choice. The on-premise model ensures that the raw audio and transcript never leave the company’s network, satisfying PCI DSS Requirement 9.4.1 for physical and logical access controls. The on-premise model also ensures that the model is not trained on the company’s data, which is a key requirement for PCI DSS compliance. The trade-off is that the on-premise model is more expensive to build and maintain, but the cost is offset by the reduction in the cost per ticket.

    The Recommendation: A 4-Week Audit and Pilot

    The recommendation is to start with a 4-week audit and pilot. The audit takes 5 business days, and the pilot takes 3 weeks. The audit includes a process mapping, a data collection, and a cost model. The pilot includes a voice agent that is built on an on-premise model and integrated with the company’s CRM, helpdesk, and Notion/Confluence. The pilot is measured against the baseline, which is the current first-response time and the current cost per ticket. The pilot is tuned based on the measurement, and the rollout is planned based on the pilot’s results. The rollout is a phased rollout, which starts with a small group of tickets and expands to the full ticket volume. The rollout is measured against the baseline, and the cost per ticket is cut by 40%.