Background: A 12-Person UAE Insurtech Preparing for Scale
This case study is a composite drawn from patterns observed across multiple engagements. No named customer is represented. The details are drawn from recurring scenarios in the field, and the metrics reflect realistic ranges rather than a single client’s exact figures.
The company in this case is a 12-person insurtech operating in Dubai, serving SMEs in the logistics and trade sectors. It writes cargo, marine, and professional liability policies. The team runs a lean stack: a custom policy management system built on PostgreSQL, a helpdesk on a mid-tier SaaS platform, and Microsoft Teams as the primary internal communication channel. The founder and two senior agents handle all customer inquiries, claims intake, and policy renewals. There is no dedicated IT team; the founder manages the stack directly. The company is in the scaling phase: it has doubled its policy book in 18 months and is preparing for a Series A raise, which requires demonstrating operational efficiency to investors.
Challenge: 4-Hour First-Response Times and a 6-Week Investor Deadline
The founder’s core complaint was not that agents were slow, but that first-response time was inconsistent and depended on which agent was on shift. The median first-response time for a policy status inquiry was 4 hours 12 minutes, but the 90th percentile exceeded 9 hours. The root cause was not agent capacity; it was that every ticket required the agent to open the policy management system, verify the policy number, check the status, and draft a response from scratch. The agent spent 11 minutes on average per ticket, and the queue grew faster than the team could clear it.
The operational pressure was twofold. First, the Series A timeline was 6 weeks out, and the investor deck needed a credible operational metric. Second, the company had just signed a new client in the logistics sector that required a 4-hour SLA on first response, which the current process could not guarantee. The founder needed a solution that could be deployed in under 3 weeks, required no new infrastructure, and kept all customer data within the UAE. GDPR compliance was not a legal requirement for a UAE-based company, but the client’s end-customers included EU-based logistics firms, and the data processing agreement required GDPR-aligned handling of personal data.
Approach: A 10-Day Build on LangGraph with a Human-in-the-Loop Gate
The engagement followed a fixed-scope pilot model. The first 3 days were a process audit: the dedicated AI team shadowed 2-3 agents, logged every ticket, and mapped the decision tree for the top 20% of ticket volume. The audit identified three ticket types that accounted for 74% of agent time: policy status inquiries, document requests (certificates of insurance, policy schedules), and simple claim status checks. These were the pilot scope. Claims adjudication, premium disputes, and health-data-related tickets were explicitly excluded.
The technical build used LangGraph to model the triage workflow as a stateful graph. The pipeline had four nodes: classify (assign ticket type and urgency), extract (pull policy number, claim reference, and document type from the ticket body), draft (generate a response using the policy management system’s API), and route (send to the appropriate agent queue with a confidence score). The model layer used the OpenAI API for classification and drafting, with a fallback to an open-weight model on the client’s own hardware for any ticket flagged as containing health data. The integration surface was the helpdesk API and Microsoft Teams: the agent received a Teams message with the AI’s draft, the extracted fields, and a one-click approve/edit/reject button. The human-in-the-loop gate was mandatory: no response went to the customer without agent approval. The entire build, including the Teams integration and the baseline measurement protocol, was completed in 10 working days. The remaining 2 days were reserved for shadowing and go-live.
Outcome: First-Response Time Down to 34 Minutes, Error Rate at 6%
The baseline was captured during the first 3 days of shadowing, before the AI was live. The median first-response time for the three in-scope ticket types was 3 hours 48 minutes. The agent time per ticket was 11.2 minutes. The error rate on manual classification (measured by comparing the agent’s routing decision against the ticket’s actual content) was 14%.
After go-live, the post-pilot measurement ran for 10 working days. The median first-response time dropped to 34 minutes. The agent time per ticket fell to 3.8 minutes, because the agent was reviewing a pre-drafted response and confirming extracted fields rather than starting from scratch. The classification error rate, measured by comparing the AI’s routing against the agent’s final decision, was 6.2%. The 90th percentile first-response time, which had been 9 hours 14 minutes, fell to 1 hour 22 minutes. The agent approval rate on AI drafts was 88%, meaning 12% of drafts required edits before approval. The most common edit was adding a policy-specific detail that the model did not have access to. No tickets involving health data or claims adjudication were processed by the AI during the pilot, as per the scope exclusion. The client reported that the 4-hour SLA for the new logistics client was met on 96% of tickets during the pilot period.
Lessons for Teams Scaling AI Across Departments
- Scope the pilot to one workflow, one channel, one integration surface. The 2-week timeline only works if the scope is narrow. Adding voice, chat, or multi-language support in the first pilot stretches the timeline and dilutes the measurement. The pilot’s job is to prove the model, not to build a platform.
- Define the approval gate before the build starts. Ambiguity about who approves what creates compliance risk and slows the go-live. In this case, the gate was clear: the agent approves, the AI drafts. For any ticket touching money, health data, or a contract, the gate is mandatory. Document the logic and retain audit logs for GDPR accountability.
- Capture the baseline before the AI is live. Without a measured before/after, the pilot cannot prove its value. The baseline should be captured during shadowing, not after go-live. Measure median first-response time, agent time per ticket, and classification error rate. The delta is the reported outcome.
- Use the messaging channel the agents already use. Integrating with Microsoft Teams or Slack means the approval workflow lives where the agent already works. A separate dashboard adds context-switching and reduces adoption. The integration should be a webhook or API call, not a custom app.
- Treat the pilot as a stepping stone, not a one-off. The pilot proves the model on one workflow. The rollout to other departments (claims, underwriting, renewals) requires a separate scope, a separate baseline, and a separate approval gate. The architecture is model-agnostic, so the same LangGraph pipeline can be extended to new workflows without a rewrite.
Leave a Reply