Deploying a RAG Assistant for Order Status Updates in Swiss E-Commerce

The Problem: Senior Support Staff Buried in Routine Order Status Tickets

Your support team at a 2,000+ employee e-commerce company in Switzerland handles thousands of order and shipment status inquiries weekly. Senior agents spend 40-60% of their time on routine lookups: “Where is my package?” “Why is my order delayed?” This work does not require judgment, but it consumes the people who should be handling complex escalations, refund disputes, and customer retention conversations. The EU AI Act, which applies to Swiss companies serving EU customers, requires transparency when AI systems interact with users. You need a retrieval-augmented knowledge assistant that drafts accurate responses from your order-management system and shipping carrier data, integrates with Zendesk or Intercom, and keeps a human in the loop for anything touching refunds or contract terms. The goal: cut first-response time from hours to minutes, reduce error rate on shipping information, and free senior staff for high-value work within a 3-month integration sprint.

Prerequisites: What You Need Before the Sprint Starts

Before starting the integration sprint, confirm these are in place:

  • Zendesk or Intercom API access: OAuth 2.0 tokens with read/write permissions for tickets, macros, and webhooks. Test with a sandbox account first.
  • Order-management system (OMS) API: Read access to order status, tracking numbers, and shipping carrier data. If you use Shopify, SAP Commerce, or a custom OMS, document the endpoint schema.
  • Shipping carrier APIs: Integration with at least your top two carriers (e.g., Swiss Post, DHL) for real-time tracking events.
  • PostgreSQL 15+ with pgvector extension: CREATE EXTENSION vector; Run on a dedicated instance with at least 16 GB RAM for 500k+ vectors.
  • LLM endpoint: OpenAI API key (gpt-4o or claude-3-5-sonnet) for drafting, or an on-prem Llama 3 70B instance if customer PII cannot leave your infrastructure.
  • EU AI Act compliance documentation: A data-protection impact assessment (GDPR Article 35) and a model card for each LLM endpoint.
  • Baseline metrics: Export 30 days of ticket data from Zendesk/Intercom. Calculate average first-response time, resolution rate, and error rate on shipping-related tickets.

Step 1: Audit the Workflow and Establish a Baseline

Run a process audit on your last 90 days of support tickets. Filter for order and shipment status inquiries: “Where is my order?” “Tracking number not working” “Delivery delayed.” Count the volume, measure average handling time, and identify the top five questions. For a 2,000+ employee e-commerce company, this typically represents 35-50% of total ticket volume. Export the data to a CSV with columns: ticket_id, subject, category, first_response_time, resolution_time, agent_id, error_flag. Calculate the baseline: if your average first-response time is 4 hours and error rate on shipping information is 8%, those are your targets to beat. Document this baseline in a one-page report. This becomes the measurement framework for the pilot and rollout phases.

Step 2: Build the RAG Pipeline with pgvector

Build the retrieval layer using pgvector. Chunk your knowledge base: shipping policies, carrier SLAs, return procedures, and order status definitions. Use a 512-token chunk size with 50-token overlap. Generate embeddings with OpenAI text-embedding-3-small (1536 dimensions) or bge-base-en-v1.5 (768 dimensions) if you prefer open-weight models. Load into PostgreSQL:

CREATE TABLE documents (
  id SERIAL PRIMARY KEY,
  content TEXT,
  metadata JSONB,
  embedding vector(1536)
);
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);

Set ef_search = 64 for sub-10 ms recall. Test with 20 sample queries: “Where is my order with tracking number XYZ?” Verify that the top-5 retrieved chunks contain the relevant shipping policy and carrier SLA. If recall is below 90%, adjust chunk size or add metadata filters (e.g., WHERE metadata->>'carrier' = 'DHL').

Step 3: Integrate with Zendesk or Intercom via Webhooks

Connect the RAG pipeline to Zendesk or Intercom. For Zendesk: create a webhook on ticket creation that triggers your RAG service. The service retrieves relevant chunks, calls the LLM endpoint with a system prompt: “You are a support assistant for [Company]. Use only the retrieved context to draft a response. If the context does not contain the answer, say so. Do not invent tracking numbers or delivery dates.” Post the drafted response to the ticket via the API with a RAG-drafted tag. For Intercom: use the Events API to trigger on ticket.created and the Agent Inbox API to post the draft. Store the correlation ID (ticket_id + timestamp) in a log table for audit trails. This satisfies EU AI Act Article 50 transparency requirements: users are informed they are interacting with an AI, and every response is traceable to its source documents.

Step 4: Add Human-in-the-Loop Approval for Sensitive Actions

Implement the human-in-the-loop approval workflow. Any RAG-drafted response that touches refunds, address changes, or contract terms must be approved by a human before sending. In Zendesk, create a custom field ai_approval_status with values: pending, approved, rejected. When the RAG service posts a draft, set ai_approval_status = pending and assign the ticket to a supervisor queue. The supervisor reviews the draft, the retrieved context, and the LLM’s confidence score. If approved, the ticket moves to approved and the response sends. If rejected, the supervisor edits or reassigns. Log every approval decision with the supervisor’s user ID and timestamp. This workflow is mandatory under EU AI Act Article 50 for any AI system that makes decisions affecting consumers. For a 3-month sprint, build a simple approval UI in React or use Zendesk’s built-in ticket views filtered by ai_approval_status = pending.

Step 5: Pilot with 10-20% of Tickets and Measure

Run the pilot with 10-20% of order-status tickets for two weeks. Route a subset of tickets (e.g., all tickets tagged order_status from a specific region or carrier) to the RAG assistant. Measure: first-response time (target: under 15 minutes vs. baseline 4 hours), resolution rate (target: 80%+ first-contact resolution), and error rate on shipping information (target: under 2% vs. baseline 8%). Sample 5% of AI-drafted responses weekly. Compare each against the OMS and carrier API data. If the assistant states a delivery date, verify it matches the carrier’s tracking event. If error rate exceeds 2%, pause the pilot, re-index the knowledge base, and adjust the LLM prompt to require citation of specific tracking events. Document every error in a log with the ticket ID, the incorrect claim, and the correct data from the OMS. This log feeds into the EU AI Act model card and the GDPR Article 35 impact assessment.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *