Background: A 24-Person Advisory Practice in Dubai
This case study is a composite drawn from patterns observed across multiple engagements. No named customer is represented. The firm, the metrics, and the timeline are representative of a recurring profile: a 20-to-30-person professional services practice in the UAE that has outgrown manual lead handling but cannot justify a dedicated sales-ops hire.
The firm in question is a 24-person advisory practice based in Dubai, serving clients across the Gulf and North Africa. Its revenue mix is 60 percent consulting, 30 percent managed services, and 10 percent training. The sales team consists of four account executives and one sales operations coordinator who also handles invoicing and reporting. The CRM is Salesforce Sales Cloud, with a custom object for engagements and a standard Lead object. Inbound leads arrive through three channels: the firm’s website form, a LinkedIn outreach sequence, and referrals from two partner firms. A significant share of inbound leads is in Arabic or French, and the sales team has historically relied on a single bilingual coordinator to translate and qualify them before an AE picks up the record.
The firm’s annual revenue is in the range of USD 3 to 5 million. It has no dedicated data team, no ML infrastructure, and no prior AI deployment. Its AI maturity, in the terms used by Forfis, is Running Isolated Pilots: the sales director has experimented with a ChatGPT prompt for drafting follow-up emails, but nothing is integrated into the CRM, and no baseline metrics exist.
Challenge: Multilingual Lead Triage Under GDPR and a Hiring Freeze
The sales director’s stated goal was simple: scale operations without new hires. The firm had just closed a USD 800,000 engagement and was onboarding two more AEs, which would push the coordinator’s workload past sustainable capacity. The coordinator was already spending roughly 12 hours per week on lead triage: reading inbound emails, translating Arabic and French summaries, assigning a priority, and updating the CRM. With two more AEs, that number would climb to 20 hours per week, effectively consuming half the coordinator’s capacity and leaving no room for the reporting and invoicing tasks that kept the finance team from chasing her for data.
The operational pressure was compounded by a GDPR and UAE data-protection constraint. The firm’s client base includes two EU-headquartered companies, and its engagement contracts require that personal data be processed under a documented lawful basis. The sales director had been told by a vendor that an AI lead-qualification tool would “just work,” but she had no clarity on where the data would be processed, who would be the data controller, or how the firm would demonstrate compliance if a client’s DPO asked for a data-flow map.
The deadline was driven by the firm’s Q3 planning cycle. The sales director needed a working pilot in the CRM before the Q3 forecast was locked, which gave an 8-week window from kickoff to a measurable baseline comparison. The budget was capped at a level that excluded a full-time data engineer hire; the solution had to be delivered as an Integration Sprint by an external product studio.
Approach: An 8-Week Integration Sprint on Salesforce
Forfis ran an 8-week Integration Sprint structured in three phases. Weeks 1 to 2 were a process audit: the Forfis team shadowed the coordinator for three days, mapped every touchpoint in the lead lifecycle, and identified the two workflows with the highest time-to-value: (1) multilingual lead translation and initial qualification, and (2) data enrichment of lead records with firmographic and engagement-history fields that the coordinator was filling manually from public sources.
Weeks 3 to 5 were the pilot build. The technical stack was the OpenAI API (GPT-4o) for classification and translation, with a thin Python service that read Lead objects from Salesforce via the REST API, called the model, and wrote the enriched fields back. The service ran on a single AWS t3.medium instance in the eu-west-1 region, with all API calls logged to an S3 bucket for audit. The human-in-the-loop gate was implemented as a Salesforce approval process: the agent wrote a draft score and rationale to a custom field, and the coordinator approved or rejected it from a standard Salesforce queue. No record was marked “Qualified” until a human clicked approve.
Weeks 6 to 8 were rollout and baseline measurement. The pilot ran on 100 percent of inbound leads for four weeks. The Forfis team tracked cycle time (timestamp from lead creation to “Qualified” status) and error rate (records where the coordinator overrode the agent’s score by more than 20 points) against the pre-pilot baseline collected during the audit.
Outcome: Cycle Time Down 87 Percent, Error Rate at 4 Percent
The pre-pilot baseline, measured over the three days of the audit, showed a median cycle time of 48 hours from lead creation to qualified status, with a 90th percentile of 96 hours. The error rate on manual qualification was not measured before the pilot, so the team established it retrospectively: during the first two weeks of the pilot, the coordinator reviewed 120 leads and flagged 14 where the agent’s score diverged from her own judgment by more than 20 points, an error rate of roughly 12 percent.
By week 8, the median cycle time had dropped to 6 hours, with the 90th percentile at 18 hours. The error rate on the agent’s scores, measured against the coordinator’s overrides, had fallen to 4 percent after the team added a 30-term glossary for Arabic business terminology (contract values, service tiers, compliance references) to the prompt. The coordinator’s weekly time spent on lead triage dropped from 12 hours to approximately 3 hours, freeing capacity for the reporting and invoicing tasks that had been slipping.
The firm did not hire a new sales-ops coordinator. The two new AEs onboarded on schedule. The sales director reported that the Q3 forecast was locked on time, and the firm’s two EU clients’ DPOs accepted the data-flow map and DPA without further questions. The pilot was extended to the French-language lead stream in week 9, and the firm is evaluating a second use case (document extraction from engagement letters) for Q4.
Lessons for Similar Teams
- The glossary is the highest-leverage artifact. The 30-term Arabic business glossary reduced the error rate from 12 to 4 percent more than any prompt engineering change. Teams in multilingual markets should budget time for a domain-specific glossary during the audit phase, not after the pilot shows errors.
- The human-in-the-loop gate is not optional in week one. The coordinator’s overrides in the first two weeks surfaced three classification errors that the model would have silently propagated. Removing the gate before the error rate is below 2 percent for two consecutive weeks is the single most common mistake Forfis sees in isolated pilots.
- The CRM API is the integration surface, not the model. The entire pilot ran on standard Salesforce REST calls. No custom middleware, no iPaaS, no new database. Teams that over-architect the integration layer burn the 8-week window on plumbing instead of on the classification logic that actually moves the metric.
- GDPR compliance is a data-mapping exercise, not a legal opinion. The firm’s DPA with OpenAI and the data-flow map were drafted in week 2, during the audit, not in week 8. Waiting until the pilot is live to address data-protection questions creates a compliance gap that is harder to close retroactively.
- The 8-week window is realistic only if the audit is front-loaded. Two weeks of shadowing and process mapping before any code is written is non-negotiable. Teams that compress the audit to three days to “save time” typically spend weeks 4 to 6 reworking the classification logic because the initial prompt was built on an incomplete understanding of the lead lifecycle.