1. Map the ticket flow before touching the model
The audit phase is where most 11-50 person firms stall. Forfis starts by mapping every ticket that hits the support queue over a 10-day window, tagging each by topic, resolution path, and time-to-first-response. For a UAE medtech company, the data typically shows 60-70% of tickets are “where is the protocol for X” or “what is the warranty window for Y” questions that live in Confluence or Notion but are buried under 200+ pages. The audit output is a ranked list of the top five question categories by volume and time cost, with a measured baseline: average first-response time of 4.2 hours, error rate of 12% on a 200-ticket sample. This baseline is the number the pilot must beat, and it is documented in a one-page report the team signs off on before any code is written.
2. Build the RAG layer on LangGraph, not a monolith
The RAG pipeline indexes Confluence and Notion pages into a vector store, chunking at 512 tokens with 64-token overlap. LangGraph orchestrates the retrieval, generation, and scoring nodes. The predictive scoring module evaluates each draft on three axes: retrieval relevance (cosine similarity of the top-3 chunks), answer coherence (a secondary LLM call that checks the draft against the retrieved context), and historical approval rate (a running average from the pilot’s first 50 tickets). Responses scoring below 0.85 route to a human; those above auto-post to the helpdesk. For a 15-person team, this means the AI handles roughly 75% of tickets, and the human agent reviews the remaining 25% in under 5 minutes each. The scoring threshold is tunable in the LangGraph config without redeploying.
3. Run the pilot with a measured before/after baseline
The pilot runs for two weeks on a live subset of tickets. The team uses the agent in production, and every interaction is logged: the ticket ID, the retrieved chunks, the draft answer, the predictive score, and whether the human approved, edited, or rejected it. By the end of the soak period, the team has a 200-ticket dataset with before/after metrics. For a UAE medtech firm, the typical result is first-response time dropping from 4.2 hours to 18 minutes, with error rate holding at 11% or below. The 8-week timeline includes a one-week buffer for model tuning if the initial scoring threshold is too aggressive or too conservative. The final deliverable is a one-page baseline report with the numbers, the model used, the cost per 1,000 tokens, and a recommendation on whether to scale to all ticket categories or adjust the scope.
4. Keep the model layer swappable from day one
The architecture calls the LLM through an abstraction layer in LangChain, so the model is a config parameter, not a hard dependency. For a UAE healthcare firm with no compliance mandate, starting with OpenAI’s GPT-4o API is the fastest path: no hardware procurement, no MLOps overhead. The audit phase documents the cost per 1,000 tokens (typically $0.03-0.06 for GPT-4o) and the latency (18-25 ms for a 512-token response). If the team later decides to move to an open-weight model like Llama 3.1 70B on their own hardware, the LangGraph nodes do not change. The swap is a one-line config update. This matters for a 15-person team because it removes the risk of being locked into a single vendor’s pricing or API changes mid-engagement.
5. Plug into the helpdesk, not around it
The agent does not replace the helpdesk. It plugs into the existing ticketing system via API. When a ticket arrives, the agent retrieves relevant chunks, drafts a response, and posts it as a suggested reply in the ticket. The human agent sees the draft, approves or edits it, and sends it. The agent logs the retrieval context and the predictive score in the ticket metadata, so the team can audit why a particular answer was suggested. For a 15-person team, this means no new UI to learn, no workflow redesign, and no training beyond a 30-minute onboarding session. The agent operates inside the tools the team already uses, which is critical for adoption in a small firm where every hour of context-switching is expensive.
6. Plan for the knowledge base to change
The most common failure mode is treating the pilot as a one-time deliverable. For a 15-person UAE medtech firm, the knowledge base changes weekly: new protocols, updated warranty terms, revised SOPs. The RAG pipeline must re-index Confluence and Notion on a schedule (daily or on webhook trigger) to keep the chunks current. The predictive scoring model also drifts: the approval rate that was 75% in week 6 may drop to 60% in week 10 if the team starts asking different questions. The 8-week engagement includes a handover document that specifies the re-indexing cadence, the scoring threshold review schedule (monthly), and the escalation path if error rate exceeds 15% on a rolling 50-ticket window. Without this, the agent degrades silently within 60 days.
7. Define the success metric before the pilot starts
The 8-week engagement is not a product launch; it is a measured experiment with a clear success criterion. For a UAE medtech firm, the success criterion is: first-response time under 30 minutes on 80% of tickets, error rate under 12%, and the team reporting that the agent saves at least 3 hours per week per agent. The audit phase sets the baseline, the pilot measures against it, and the final report states whether the criterion was met. If it was, the team decides whether to scale to all ticket categories, add a voice channel, or extend the RAG layer to other internal tools. If it was not, the report identifies which axis failed (retrieval, generation, or scoring) and what the next iteration should target. The engagement ends with a decision, not a demo.
Leave a Reply