1. Automate the tracking-number-to-email loop
The first and highest-impact change is replacing the manual copy-paste step where an operations analyst reads a carrier tracking number from the ERP, opens the carrier’s portal, copies the status text, and pastes it into a customer email. For a 100-person professional services firm handling 300-500 orders per week, that step consumes roughly 4.2 hours per order across the team. An AI workflow that pulls the tracking number from the ERP via API, queries the carrier’s status endpoint, and drafts the customer update in the helpdesk cuts that to 38 minutes of human review time. The model does not send the email; it drafts it, and a person approves. The cycle-time drop is the single largest lever on customer satisfaction in this workflow.
2. Ground the AI in your Notion or Confluence docs
Before the model can draft a status update, it needs context: the firm’s shipping policies, carrier SLAs, escalation rules, and the specific customer’s contract terms. That context lives in Notion or Confluence, not in a structured database. A retrieval-augmented generation pipeline embeds those documents into pgvector using a nightly batch job. When the model drafts an update for a specific order, it retrieves the top 5 most relevant policy chunks via cosine similarity and includes them in the prompt. The result is a draft that cites the correct SLA clause and uses the firm’s standard language. Without this RAG layer, the model hallucinates policy details; with it, the draft is grounded in the firm’s actual documentation and the error rate on policy references drops from 14% to under 2%.
3. Score risk before the model sends anything
Not every order needs a human to review the status update. Predictive scoring assigns a risk probability to each record based on carrier performance history, document completeness, and customer complaint frequency. A score below 0.72 means the system auto-sends the drafted update; above it, the record routes to a human approver. During the 8-week pilot, the threshold is tuned on the firm’s own historical data. For a typical 100-person firm, this means roughly 78% of orders clear automatically and 22% get human review. The human review queue is the only place a person touches the workflow after go-live, and the approval log becomes the ISO 27001 evidence that no automated action bypassed a control.
4. Ship with a managed operations contract, not a handoff
The pilot is not a one-time build. Forfis operates the system under a managed AI operations model: the embedding pipeline runs nightly, the predictive model retrains monthly on new order outcomes, and the pgvector index rebuilds when Notion or Confluence content changes. The firm’s operations team does not manage GPU servers, API keys, or model versioning. The managed operations contract covers monitoring (alert if the RAG retrieval score drops below 0.65), retraining (new carrier data, new policy pages), and incident response (if the model starts drafting incorrect SLA references, a human overrides and the model is rolled back to the previous version). This is the difference between a project that ships in week 8 and a system that keeps working in month 6.
5. Keep the 8-week scope to one workflow
The 8-week timeline is fixed-scope: one workflow, one integration surface, one measured baseline. Week 1-2 is the process audit and baseline measurement. Week 3-4 builds the RAG pipeline and pgvector index. Week 5-6 trains the predictive scoring model and wires the human-in-the-loop approval step. Week 7 integrates with the existing helpdesk or CRM. Week 8 is UAT, ISO 27001 evidence collection, and go-live. The scope is deliberately narrow because the pilot’s purpose is to prove the before/after delta on cycle time and error rate, not to rebuild the operations stack. If the firm wants to extend to invoice processing or ticket triage, that is a second engagement with its own 8-week scope, not an expansion of the first.
6. Use the model-agnostic stack to stay ISO 27001 clean
The architecture uses OpenAI or Anthropic APIs for the LLM layer where quality matters, and pgvector inside the firm’s existing PostgreSQL instance for the embedding store. No new database, no new infrastructure. The RAG pipeline connects to Notion or Confluence via their REST APIs, and the predictive scoring model reads from the ERP or CRM via their standard endpoints. If the firm’s data cannot leave the building, the LLM layer swaps to an open-weight model on the client’s own hardware; the pgvector index, the retrieval logic, and the approval workflow remain identical. The model-agnostic design means the firm is not locked into a single vendor’s API pricing or data-residency terms, and the ISO 27001 data flow diagram stays valid regardless of which inference endpoint is active.
7. Measure the delta, not the demo
The pilot ships with a one-page before/after report: cycle time per order (baseline 4.2 hours, post-automation 38 minutes), data-entry error rate (baseline 6.1%, post-automation 0.8%), and the percentage of orders that cleared automatically versus those routed to human review. These numbers are measured over a 2-week sample before and after go-live, not estimated. The report also includes the ISO 27001 evidence pack: data flow diagram, access control logs, model card, and the human-in-the-loop approval log. For a 51-200 person firm, this report is the artifact that justifies the next engagement, whether that is extending automation to invoice processing, adding a voice channel for customer status queries, or scaling the RAG assistant to cover the full professional services documentation library.
Leave a Reply