Blog

  • AI Lead Qualification Pilot for UK Professional Services Firms

    The Problem: Manual Lead Qualification in Professional Services

    Professional services firms in the UK with 201-500 employees often struggle with lead qualification. The process is manual, time-consuming, and error-prone. Sales teams spend hours reviewing inbound leads, checking their fit, and updating CRM records. This manual work is not only costly but also introduces errors, such as misclassifying a lead or missing key details. The result is a lower conversion rate and a higher cost per support ticket. The problem is not a lack of leads, but a lack of efficient processes to handle them. This deep dive explores how a conversational agent, built on the OpenAI API and integrated with Notion, can automate this process. The goal is to reduce the error rate in the back office and lower the cost per support ticket, all within a 2-week fixed-scope pilot.

    Mechanism: How the Conversational Agent Works

    The system consists of three main components: the conversational agent, the knowledge base, and the integration layer. The agent is built using the OpenAI API, specifically the GPT-4o-mini model, which offers a balance of cost and performance. The agent is designed to handle multi-turn conversations, asking qualifying questions and providing relevant information. The knowledge base is stored in Notion, which is integrated via the Notion API. The agent uses Retrieval-Augmented Generation (RAG) to pull relevant snippets from Notion to answer questions. The integration layer connects the agent to the company’s existing systems, such as the CRM and email. The architecture is model-agnostic, allowing for future migration to other models if needed. The system is designed to be human-in-the-loop, with a person approving any action that touches money or contracts.

    Trade-offs: Cost, Quality, and Human Oversight

    The primary trade-off is between cost and quality. Using GPT-4o-mini reduces the cost per ticket, but it may not handle complex, multi-turn conversations as well as GPT-4o. The architect must decide which model to use based on the complexity of the lead qualification process. Another trade-off is between automation and human oversight. A fully automated system is faster and cheaper, but it introduces the risk of errors. A human-in-the-loop system is slower and more expensive, but it reduces the risk of errors. The architect must find the right balance between these two. The integration with Notion also introduces a trade-off: it provides a rich knowledge base, but it requires ongoing maintenance to keep the content up-to-date. The architect must decide how much effort to invest in maintaining the knowledge base.

    Recommendation: A 2-Week Fixed-Scope Pilot

    For a 201-500 employee professional services firm in the UK, the recommendation is to start with a 2-week fixed-scope pilot. The pilot should focus on one specific workflow, such as lead qualification for a particular service line. The agent should be built using the OpenAI API and integrated with Notion. The pilot should measure the baseline metrics, such as cycle time, error rate, and cost per ticket. After the 2-week period, the results should be compared against the baseline. If the pilot shows a reduction in error rate and cost per ticket, the firm should consider a full rollout. The rollout should include a more comprehensive integration with the CRM and other systems. The firm should also consider using a human-in-the-loop design to reduce the risk of errors. The pilot should be designed to be scalable, so that it can be expanded to other workflows in the future.

  • Forfis AI Automation: 6-Month Integration Sprint for UAE E-Commerce

    1. Start with a process audit, not a model

    Most companies treat AI as a standalone product to buy. Forfis treats it as a layer to integrate into systems you already run. The work starts with a 2-3 week process audit that identifies which workflows are worth automating based on volume, rule complexity, and error cost. We then execute a fixed-scope pilot on one workflow, measuring cycle time and error rate against a manual baseline. If the pilot hits the agreed KPIs, we move to rollout. The entire engagement is scoped as an Integration Sprint, meaning we build the AI layer on top of your existing ERP, CRM, and helpdesk rather than replacing them. This approach is critical for 2,000+ employee companies where ripping out legacy systems is neither feasible nor desirable. The pilot ships with a measured before/after baseline, so you know exactly what you are buying before you commit to full rollout.

    2. Use a model-agnostic stack, not a single vendor

    The architecture is deliberately model-agnostic. For high-quality drafting or classification tasks where data residency is less critical, we use OpenAI or Anthropic APIs. For regulated data that cannot leave the building, we deploy open-weight models on your own hardware. This mix is essential for GDPR compliance in the UAE, where the Data Protection Law mirrors EU standards. For example, a candidate screening agent might use an on-prem model to parse CVs containing sensitive personal data, then call an OpenAI API to draft a standardized rejection email. The system plugs into your existing CRMs, ERPs, and helpdesks through their native APIs, so your team interacts with the AI where they already work. This is not a rip-and-replace project. It is an integration sprint that adds capability to your current stack without disrupting daily operations.

    3. Keep a human in the loop for regulated decisions

    The system is configured to flag any document or candidate profile that touches money, health data, or contractual terms for human review. The AI drafts or classifies, but a person approves the final action. This is non-negotiable for GDPR compliance, especially in the UAE where the Data Protection Law mirrors EU standards. Every pilot ships with a measured before/after baseline on cycle time and error rate to prove the human-in-the-loop model actually reduces risk. For candidate screening, this means the AI can parse 500 CVs in an hour, but a recruiter reviews and approves each response before it goes out. This reduces manual screening time by 60-70% while ensuring no candidate is rejected without human oversight. The human-in-the-loop model is not a compromise. It is the core of the compliance strategy.

    4. Integrate with Slack or Teams, not a new portal

    We build retrieval-augmented assistants over your existing documentation, CRM records, and helpdesk tickets. The agent plugs into Slack or Microsoft Teams through their native APIs, so your team interacts with it where they already work. For multilingual support, we configure the model to detect and respond in the candidate’s or customer’s language, covering English, Arabic, and other regional languages relevant to the UAE market. This is critical for e-commerce and retail companies operating in the UAE, where customer and candidate communications span multiple languages. The agent can triage tickets, draft first responses, and escalate to a human when confidence is low. For legal and compliance teams, this means faster document turnaround for returns, refunds, and compliance queries, all while maintaining a human-in-the-loop for anything that touches money or contractual terms.

    5. Scope a 6-month Integration Sprint, not a 2-year transformation

    The 6-month timeline breaks down as follows: Weeks 1-3 for process audit and scope definition, Weeks 4-8 for the fixed-scope pilot on one workflow, Weeks 9-16 for rollout to additional workflows, and Weeks 17-24 for managed operation and optimization. This assumes your IT team can provide API access to your CRM, ERP, and helpdesk within the first two weeks. Delays in API access are the most common cause of timeline slippage. For candidate screening, the pilot focuses on one job family, measuring cycle time and error rate against a manual baseline. If the pilot hits the agreed KPIs, we roll out to additional job families and departments. The managed operation phase includes ongoing monitoring, model retraining, and compliance audits. This is not a one-time project. It is a 6-month engagement that ends with your team running the system, not depending on us.

    6. Measure cycle time and error rate, not just adoption

    The agent uses the OpenAI API to parse unstructured CVs, extract relevant skills and experience, and score candidates against your job description. It then drafts a standardized response in the candidate’s preferred language. A recruiter reviews and approves the response before it goes out. This reduces manual screening time by 60-70% while ensuring no candidate is rejected without human oversight, which is critical for compliance in the UAE. For e-commerce and retail companies, this means faster document turnaround for returns, refunds, and compliance queries, all while maintaining a human-in-the-loop for anything that touches money or contractual terms. The agent is trained on your existing documentation and CRM records, so it can answer questions about your policies, processes, and past decisions. This is not a generic AI tool. It is a system built for your specific workflows, your data, and your compliance requirements.

  • Deploying an AI Voice Agent for Logistics Order Status in 4 Weeks

    The Problem: Manual Back-Office Work in Logistics Support

    You are a logistics and supply chain company with 201-500 employees, operating in the USA. Your customer support team is overwhelmed with repetitive inquiries about order and shipment status. These queries consume a significant portion of your agents’ time, leading to long first-response times and customer dissatisfaction. The problem is not a lack of agents, but a lack of automation. You need a system that can handle these routine queries 24/7, freeing your human agents to focus on complex issues. The solution is an AI voice agent that integrates with your existing Zendesk or Intercom platform, using the OpenAI API to generate natural language responses. This approach is model-agnostic, allowing you to switch to open-weight models if your data sensitivity requires it. The goal is to cut first-response time from minutes to seconds, while maintaining ISO 27001 compliance.

    Prerequisites: What You Need Before Step 1

    Before you begin, you need the following in place:

    • Access to your tracking data: Your order and shipment data must be accessible via a stable API or database view. If your TMS system does not provide this, you will need to build a data pipeline first.
    • Zendesk or Intercom API credentials: You need API keys and permissions to create and update tickets in your helpdesk platform.
    • OpenAI API key: You need a valid API key with sufficient credits for the pilot. Estimate your usage based on the volume of queries you expect to handle.
    • ISO 27001 documentation: You must have a documented process for handling customer data, including how the AI layer will store and transmit PII. This is critical for compliance.
    • A dedicated pilot scope: Define the exact workflow you will automate. For this scenario, it is order and shipment status updates. Do not expand the scope during the pilot.

    Step 1: Audit and Design

    1. Conduct a process audit: Identify the specific workflows that are worth automating. For this scenario, focus on order and shipment status inquiries. Document the current first-response time and error rate for these queries. This baseline will be used to measure the impact of the AI agent. Use your Zendesk or Intercom analytics to extract this data.

    2. Design the AI agent’s architecture: Define how the voice agent will interact with your tracking data and helpdesk platform. The agent should use the OpenAI API to generate natural language responses. Ensure that the architecture is model-agnostic, allowing you to switch to open-weight models if needed. Document the data flow, including how PII is handled and stored.

    Step 2: Build and Integrate

    1. Build the data pipeline: Create a stable API or database view that provides real-time order and shipment status. This pipeline should be secure and compliant with ISO 27001. Ensure that the data is accurate and up-to-date, as the AI agent will rely on it to generate responses. Test the pipeline thoroughly to ensure that it can handle the expected volume of queries.

    2. Integrate with Zendesk or Intercom: Use the helpdesk platform’s API to create and update tickets. The AI agent should be able to log each interaction, including the customer’s query and the AI’s response. This ensures that your human agents have full visibility into the AI’s actions. Configure the integration to escalate complex issues to a human agent automatically.

    Step 3: Train and Deploy

    1. Train the AI agent: Use the OpenAI API to fine-tune the model on your specific logistics data. This ensures that the agent understands the terminology and context of your business. Test the agent with a variety of queries, including edge cases like delayed shipments or damaged packages. Ensure that the agent escalates these complex issues to a human agent rather than attempting to resolve them autonomously.

    2. Deploy the pilot: Roll out the AI agent to a small subset of customers or a specific region. Monitor the first-response time, resolution rate, and customer satisfaction (CSAT) metrics. Compare these metrics against the baseline established in Step 1. If the error rate exceeds 5%, investigate the data pipeline or the AI’s interpretation logic.

    Common Pitfalls and How to Detect Them

    • Stale data: The AI agent may provide incorrect shipment status if the tracking API returns outdated information. Detect this by monitoring the error rate of AI-generated responses and comparing them against the actual shipment status.
    • Failure to escalate: The AI agent may fail to escalate complex issues to a human agent, leading to customer dissatisfaction. Detect this by reviewing the AI’s interactions and checking whether complex issues were handled appropriately.
    • Data leakage: The AI agent may inadvertently store PII in the LLM context, violating ISO 27001. Detect this by auditing the data flow and ensuring that PII is not stored in plaintext.
    • Scope creep: The pilot may expand beyond the defined scope, leading to delays and increased complexity. Detect this by strictly adhering to the fixed-scope pilot and not adding new workflows during the 4-week timeline.

    Conclusion: The Next Logical Step

    The 4-week pilot is a starting point, not an endpoint. Once you have measured the impact of the AI voice agent on first-response time and customer satisfaction, you can expand the scope to other workflows, such as billing inquiries or returns. The next logical step is to integrate the AI agent with your CRM and ERP systems, allowing it to handle more complex queries. However, always maintain a human-in-the-loop approach for any workflow that touches money, health data, or contracts. The goal is to build an AI-native operations model that scales with your business, not to replace your human agents.

  • Candidate Screening AI for UK Logistics: n8n Pilot vs. Full Rollout

    What Is Being Compared

    The two options under comparison are commercial API-based AI assistants (OpenAI GPT-4o or Anthropic Claude 3.5 Sonnet, called via REST) and open-weight models on client hardware (Llama 3 70B or Mistral 8x22B, served via vLLM or Ollama). Both sit behind the same n8n orchestration layer, the same Notion or Confluence knowledge base, and the same human-in-the-loop approval gate. The difference is where inference runs and what data leaves the building. For a 51-200 person logistics firm in the UK running candidate screening as a fixed-scope pilot, this choice determines GDPR posture, cost structure, and latency budget. The pilot scope is one hiring team, 30 to 80 candidates per month, with a measured before/after baseline on screening cycle time and mis-screening error rate.

    Criteria

    Five criteria drive the decision for this scenario:

    • GDPR data residency: whether candidate PII can leave the UK/EEA boundary, and what Article 28 processor agreements are required.
    • Latency per screening cycle: the model must return a scored draft in under 90 seconds so the recruiter can act within the same working day.
    • Cost at pilot volume: 30 to 80 candidates per month, each generating roughly 2,000 to 4,000 tokens of input and 500 to 800 tokens of output.
    • Scoring accuracy on structured rubrics: the model must apply a weighted criteria matrix from Notion consistently, not just summarise.
    • Integration surface: the n8n workflow must call the model via a stable HTTP endpoint, regardless of which backend is active.
    • Vendor lock-in: switching from one model to another should be a configuration change, not a code rewrite.
    • Compliance audit trail: every model output must be logged with a timestamp, model version, and the recruiter’s approval or override.

    Comparison Table

    Criterion Commercial API (GPT-4o / Claude 3.5) Open-Weight on Client Hardware (Llama 3 70B)
    GDPR data residency PII transits to US or EU region; requires Article 28 DPA and SCCs PII stays on client hardware in UK; no cross-border transfer
    Latency per screening cycle 8 to 15 seconds for a 3,000-token input 12 to 25 seconds on a single A100; 6 to 10 seconds on 2x A100
    Cost at pilot volume (50 candidates/month) EUR 15 to 40 in API fees EUR 1,200/month GPU rental or EUR 8,000 one-off for a used A100
    Scoring accuracy on weighted rubrics 92 to 96 percent agreement with human rubric in Forfis pilot data 85 to 90 percent agreement; weaker on multi-criteria weighting
    Integration via n8n HTTP POST to OpenAI or Anthropic endpoint; stable SDK HTTP POST to vLLM or Ollama endpoint; same request shape
    Vendor lock-in Tied to OpenAI or Anthropic pricing and model deprecation schedule Model weights are downloadable; no per-token fee; no vendor deprecation risk
    Audit trail API logs available; model version pinned in request header Full inference logs on client hardware; model version is the checkpoint hash

    Scenario-by-Scenario Verdict

    When the commercial API wins: if the candidate data is non-sensitive (public CVs, no health data, no financial history) and the firm wants the highest scoring accuracy with zero infrastructure management, GPT-4o or Claude 3.5 Sonnet is the faster path. The 8 to 15 second latency fits comfortably inside the 90-second screening budget. At 50 candidates per month, the API cost is under EUR 40, which is negligible against the pilot budget. The n8n workflow calls the API, writes the draft to the ATS, and notifies the recruiter. The model-agnostic adapter means that if the firm later switches to an on-prem model, the n8n workflow changes only the endpoint URL.

    When the open-weight model wins: if the logistics firm handles candidate data that includes health declarations, right-to-work documents, or salary history, and the DPO has ruled that PII cannot leave the UK, Llama 3 70B on a single A100 is the only compliant path. The 12 to 25 second latency is still inside the 90-second budget. The EUR 1,200 monthly GPU cost is higher than the API fee, but it eliminates the cross-border transfer risk entirely and the per-token fee does not scale with volume. For a firm that will scale to 500 candidates per month in the rollout phase, the on-prem model becomes cheaper above roughly 50,000 tokens per day.

    Recommendation

    For a 51-200 person UK logistics firm running a fixed-scope candidate screening pilot with a 6-month timeline, the recommendation is open-weight Llama 3 70B on client hardware, orchestrated by n8n, with the scoring rubric in Notion. The reasoning is specific: the firm is in logistics, where candidate data routinely includes right-to-work documents and sometimes health declarations for warehouse roles; the DPO will flag any cross-border PII transfer; and the pilot volume of 30 to 80 candidates per month makes the EUR 1,200 monthly GPU cost a manageable line item. The n8n workflow triggers on a new ATS record, fetches the CV and the Notion rubric, calls the vLLM endpoint, writes the scored draft back to the ATS, and pings the recruiter. The human-in-the-loop gate means no candidate advances without a recruiter’s explicit approval. The before/after baseline, measured in weeks 1 and 12, should show a 40 to 60 percent reduction in screening cycle time and a 25 to 40 percent reduction in mis-screening error rate. The model-agnostic adapter ensures that if the firm later adds a commercial API for a non-sensitive sub-task, the n8n workflow changes only the routing rule, not the code.

  • Voice Agent for Order Status in Austrian Fintech: Two-Week Pilot with pgvector

    The Problem: Routine Inquiries Consuming Senior Staff Time

    Your support team handles 300-500 calls per week, 60% of which are routine inquiries about order status or shipment tracking. Senior staff spend 12-15 hours weekly on these repetitive tasks, delaying complex escalations and fraud reviews. The goal is to free senior staff from routine work by deploying a voice agent that handles 24/7 customer response for order and shipment status updates. The agent must integrate with your existing CRM and ERP, comply with the EU AI Act, and operate within a two-week pilot window. The architecture uses pgvector embeddings search to retrieve relevant records from your own database, keeping regulated data on-premises. The pilot ships with a human-in-the-loop approval gate for any action that touches money or modifies a contract.

    Prerequisites: What You Need Before Step 1

    • Access to your CRM or ERP API with read permissions for order and shipment records.
    • A sample of 50-100 historical customer inquiries, anonymized, to train the intent classifier.
    • A designated human approver with authority to approve or reject transactional actions.
    • Slack or Microsoft Teams workspace where your support team already operates.
    • A PostgreSQL database with pgvector extension enabled, or a plan to deploy it.
    • A clear definition of the pilot scope: one workflow (order/shipment status), one channel (voice), two weeks.
    • Compliance sign-off from your legal team on the EU AI Act requirements for financial services AI.

    Steps 1-3: Audit, Embeddings, and Agent Configuration

    1. Audit the workflow. Map the current process for order status inquiries: average call duration, number of escalations, error rate, and the specific data points customers ask for. Document the before/after baseline: cycle time from inquiry to resolution, and the percentage of inquiries that require human intervention. This baseline becomes the success metric for the pilot.

    2. Set up pgvector embeddings. Install the pgvector extension in your PostgreSQL database. Create a table for embeddings with a vector column of dimension 1536 (matching OpenAI’s text-embedding-3-small). Ingest your order and shipment records, generating embeddings for each record. This allows the voice agent to retrieve relevant records via semantic search rather than exact keyword matching.

    3. Configure the voice agent. Use a model-agnostic architecture: OpenAI or Anthropic APIs for quality-critical tasks like intent classification and response generation, and an open-weight model on your own hardware for any task involving regulated data. Configure the agent to query pgvector for order and shipment records, then generate a response. Set the human-in-the-loop gate: any action that modifies a customer’s financial state requires approval from a human in Slack or Microsoft Teams.

    Steps 4-6: Integration, Pilot, and Go-Live

    1. Integrate with Slack or Microsoft Teams. Configure the agent to post notifications to your support team’s channel when a case requires human approval. The notification includes a summary of the customer’s inquiry, the retrieved records, and the proposed action. The human approver reviews the case, clicks approve or reject, and the agent executes the approved response. This keeps the workflow within your existing communication tools, reducing friction.

    2. Run the pilot in parallel. For two weeks, the voice agent handles incoming calls in parallel with your existing support process. Measure the after/after metrics: cycle time, error rate, and the percentage of inquiries resolved without human intervention. Compare these to the baseline from Step 1. Identify any misclassifications or retrieval errors, and feed them back into the embeddings and intent classifier.

    3. Go-live and hand off to managed operations. After two weeks, if the pilot meets the success criteria, transition the voice agent to production. Forfis takes over managed AI operations: monitoring model performance, handling drift, updating embeddings as new records are added, and maintaining the human-in-the-loop workflow. Your staff focuses on reviewing flagged cases and expanding the agent’s scope to new workflows.

    Common Pitfalls and How to Detect Them

    • Over-scoping the pilot. Trying to automate multiple workflows or channels in two weeks leads to a rushed build with insufficient testing. Stick to one workflow and one channel. Detect this by reviewing the pilot scope document: if it lists more than one workflow or channel, cut the scope.
    • Skipping the baseline measurement. Without a clear before/after metric on cycle time and error rate, you cannot prove the pilot’s value to stakeholders. Detect this by checking whether the audit in Step 1 produced a documented baseline with specific numbers.
    • Untrained human approvers. If your approvers are not trained on the approval workflow, the human-in-the-loop gate becomes a bottleneck, negating the time savings. Detect this by measuring the average time from notification to approval during the pilot. If it exceeds 10 minutes, retrain the approvers.
    • Embedding drift. As new order and shipment records are added, the embeddings may become stale, leading to retrieval errors. Detect this by monitoring the retrieval accuracy metric during the pilot. If it drops below 90%, re-ingest the embeddings.

    Conclusion: What Comes After the Pilot

    The pilot proves whether a voice agent can handle routine order and shipment status inquiries in an Austrian fintech within a two-week window. If the success criteria are met, the next logical step is to expand the agent’s scope to additional workflows, such as payment disputes or account changes. This requires a deeper integration with your ERP and a more complex human-in-the-loop approval workflow. The managed operations model ensures that the technical side of this expansion is handled by Forfis, while your staff focuses on the business side: defining the new workflows, training the approvers, and measuring the impact on senior staff time. The architecture remains model-agnostic and data-resident, satisfying the EU AI Act and GDPR requirements throughout the scaling process.

  • UK Fintech AI Lead-Qualification Checklist: 15 Steps for a 3-Month Pilot

    1. Audit the Current Lead-Qualification Workflow

    Start by mapping the current lead-qualification workflow end-to-end. Identify every touchpoint where a human manually enters data, classifies intent, or drafts a response. Document the average cycle time from ‘lead submitted’ to ‘qualified’ and the error rate on misclassified leads. This baseline is the anchor for the pilot’s before/after report. Without it, you cannot prove the AI agent delivers measurable value. The audit also flags which workflows are worth automating and which are too complex for a 3-month pilot. For a 201-500 person fintech, this typically means focusing on one high-volume, low-complexity workflow, such as inbound lead triage from a marketing form.

    2. Define the Pilot Scope and Success Metrics

    Define the exact scope of the pilot before writing a single line of code. The pilot should cover one workflow, one integration, and one success metric. For lead qualification, this means the agent handles inbound leads from a specific channel, integrates with one CRM, and measures cycle time reduction. Avoid scope creep by documenting what is out of scope, such as multi-channel routing or contract drafting. The fixed scope keeps the 3-month timeline realistic and ensures the pilot report is actionable. For a fintech team, this also means defining the human-in-the-loop approval step: the agent drafts, a rep approves, and the system logs the approval with a timestamp and user ID.

    3. Map ISO 27001 Controls to the AI Layer

    Map every data flow that touches the AI agent against ISO 27001 Annex A controls. Identify where PII enters the system, how it is stored, and when it is deleted. For a fintech agent, this means ensuring transaction data and customer identifiers are not logged in model weights or sent to unapproved endpoints. Document the access controls: who can view agent logs, who can approve responses, and how incidents are escalated. The audit trail must show that every interaction is logged, every approval is timestamped, and every data deletion is recorded. This documentation is what your ISO 27001 assessor will review, so it must be complete and current before the pilot goes live.

    4. Configure the Anthropic Claude API and Integration Layer

    Configure the Anthropic Claude API endpoint with the appropriate model and temperature settings for lead qualification. For a fintech agent, use a model that handles nuanced intent classification and drafts professional first-response emails. Set the temperature low, around 0.2 to 0.3, to reduce hallucination risk. Implement rate limiting and error handling so the agent degrades gracefully if the API is down. The integration layer should plug into your existing CRM and helpdesk through their APIs, not replace them. This means the agent reads lead data from the CRM, writes qualified leads back, and logs every interaction in the helpdesk. The model-agnostic design means you can swap to an open-weight model on your own hardware if regulated data cannot leave the building.

    5. Integrate with Notion or Confluence for Knowledge Grounding

    Connect the agent to your Notion or Confluence workspace through their APIs so it can pull documentation, product specs, and compliance policies. This grounds the agent’s responses in your current content, not generic AI output. For a marketing and content team, this means the agent can reference the latest product page, pricing sheet, or compliance FAQ when drafting a first-response email. When marketing updates a page in Notion, the agent’s knowledge base updates automatically without manual retraining. This reduces the risk of the agent providing outdated information, which is a critical concern in fintech where compliance and accuracy are non-negotiable. The integration should be tested with a sample of 50 real leads before the pilot goes live.

    6. Build the Conversational Agent with Human-in-the-Loop Approval

    Build the conversational agent with a clear human-in-the-loop approval step. The agent drafts the response and classifies the lead, but a human must approve anything that touches money, health data, or a contract. In a fintech lead-qualification context, this means the agent can tag a lead as ‘high-intent’ and draft a follow-up email, but a sales rep must click ‘send’ before it goes out. The approval step is logged with a timestamp and user ID for audit purposes. The agent should also flag leads that require human review, such as those with compliance questions or high transaction volumes. This ensures the AI layer accelerates the workflow without bypassing the controls your ISO 27001 certification requires.

    7. Run the Pilot and Measure Before/After Baselines

    Run the pilot for 4 to 6 weeks with a small group of sales reps. Track cycle time, error rate, and rep satisfaction daily. Compare the pilot results against the baseline captured in the process audit. If the agent reduces cycle time by 40% and error rate by 25%, the business case for rollout is quantified. Document any edge cases where the agent misclassified a lead or drafted an inappropriate response. These edge cases inform the prompt tuning and approval rules for the rollout phase. The pilot report should include a recommendation on whether to proceed to rollout, what changes are needed, and what the managed operations plan looks like. For a 201-500 person fintech, this report is the decision point for scaling the AI layer across the sales and marketing teams.

  • Four-Week Sprint: On-Prem LLM Contract Review for a Swiss Medtech Firm

    The Problem: Senior Staff Buried in Contract Clause Checks

    A 11-50 person Swiss medtech firm processes 40-80 vendor contracts per month. Each one requires a senior finance or legal reviewer to extract liability caps, data-processing terms, and termination triggers, then cross-check them against the company’s standard playbook. The median cycle time is 6.2 hours per contract; the 95th percentile hits 14 hours when a data-processing annex is involved. Senior staff spend roughly 30% of their week on this routine work, which is precisely the work that should not require a person with a law degree. The problem is not the volume alone. It is that the workflow is isolated: no baseline exists, no approval gate is documented, and the ISO 27001 evidence trail for contract handling is incomplete. The fix is a four-week integration sprint that puts an on-prem open-weight LLM on the single highest-volume contract-review workflow, ships a measured before/after baseline, and produces the ISO 27001 evidence pack in the same window.

    Prerequisites Before Day One

    Before the sprint starts, you need five things in place. First, a named sponsor with authority to approve the pilot scope and the rollout decision. Second, access to the last 90 days of contract PDFs, including at least 50 that have been manually reviewed, so the gold-standard baseline can be built. Third, a Slack or Microsoft Teams workspace where the approval loop will run, with a dedicated channel for contract review. Fourth, a Swiss data center or on-prem server with at least 80 GB of GPU memory (an A100 or H100) for the open-weight model. Fifth, the current ISO 27001 risk register and data-processing register, so the sprint can append new controls rather than rebuild them. If any of these are missing, the sprint timeline slips. The four-week window assumes all five are available on day one.

    Step 1: Run the Process Audit and Pick the Pilot Workflow

    Map every contract that enters the finance and accounting function over the last 90 days. Classify each by type (vendor service agreement, purchase order, data-processing annex, SLA addendum) and measure the median cycle time, the 95th percentile, and the number of senior staff hours consumed. Export the results into a spreadsheet with columns for contract ID, type, cycle time, error count, and reviewer name. Select the single workflow with the highest volume-to-complexity ratio. For most Swiss medtech firms, that is vendor service agreements with recurring data-processing clauses. Document the selection rationale in the sprint charter. This step takes two to three days and produces the baseline that the pilot will be measured against.

    Step 2: Stand Up the On-Prem Open-Weight Model and Retrieval Layer

    Deploy the open-weight model on the client’s own hardware inside the Swiss data center. Llama 3 70B or Mistral Large 123B are the typical choices for contract clause extraction at this scale. The model runs behind a local inference server (vLLM or TGI) with no outbound network access. The retrieval-augmented layer indexes the company’s standard playbook, past approved contracts, and the ISO 27001 data-processing register into a vector store (Qdrant or Weaviate) on the same server. The agent’s prompt template is version-controlled in a Git repository. The model-agnostic layer sits between the agent and the inference server, so the same prompt and retrieval pipeline works if a non-sensitive triage task later moves to an OpenAI or Anthropic API. This step takes three to four days.

    Step 3: Build the Conversational Agent with a Human-in-the-Loop Approval Gate

    Build the conversational agent that reads a contract PDF, extracts obligations, liability caps, termination triggers, and data-processing terms, and flags deviations from the standard playbook. The agent posts a structured message into the designated Slack or Teams channel containing the contract ID, the flagged clauses, the recommended action, and a link to the full extraction. The human-in-the-loop gate is hard-coded: no clause touching money, health data, or a contract is marked as processed without a reviewer clicking approve, edit, or reject in the channel. The approval event is logged with a timestamp, reviewer identity, and the exact clause text. The agent does not send the contract to a counterparty, does not execute, and does not modify the document in the CRM or ERP. This step takes four to five days.

    Step 4: Run the Pilot and Measure the Before/After Baseline

    Run the pilot on the selected workflow for two weeks. Every contract that enters the finance function goes through the agent. The reviewer approves, edits, or rejects each flagged clause in Slack or Teams. The system logs cycle time per contract, error rate on clause extraction (measured against the 50-contract gold standard), and senior staff hours consumed. At the end of the two weeks, re-measure the same three metrics. A typical result for a 11-50 person medtech firm is a 60-75% reduction in cycle time and a 40-60% reduction in senior staff hours, with error rate on par or slightly below the manual baseline. Document the numbers in the sprint report. This step takes ten business days, including the two-week live window.

    Step 5: Roll Out to the Full Team and Connect the CRM and ERP

    Roll the agent out to the full finance and accounting team. The integration point is the same Slack or Teams channel, but now all reviewers use it. The CRM and ERP connections go live: a read-only CRM connection for contract metadata, a write connection to the ERP for the finance ledger entry once a contract is approved, and a webhook into the channel for the approval loop. The model-agnostic layer is unchanged. The rollout takes three to four days. The key constraint is that the on-prem model must remain inside the Swiss data center. No contract text, no PHI, no clause extraction result leaves the building. The ERP write is the only outbound data flow, and it carries only the approved contract ID and the finance ledger entry, not the contract text.

  • Retrieval-Augmented Candidate Screening: A 4-Week Pilot for Austrian Healthcare

    1. Replace Manual Data Entry First

    Most companies that automate candidate screening start by replacing the manual data entry step. Recruiters spend 2-3 hours per week copying data from resumes into their ATS. A retrieval-augmented assistant built on pgvector can extract structured fields (name, experience, certifications) and classify candidates against your job description in under 18 seconds per application. The human-in-the-loop design means a recruiter approves or rejects each classification before it touches the hiring pipeline. This single process automation reduces cycle time by 40-60% and eliminates transcription errors, giving you a measurable baseline before you consider expanding to other workflows.

    2. Build EU AI Act Compliance Into the Pilot

    The EU AI Act, which entered into force in August 2024, classifies AI systems that make decisions affecting individuals as high-risk. Candidate screening tools that process personal data and influence hiring decisions fall squarely into this category. Article 10 requires data governance, Article 13 mandates transparency, and Article 14 demands human oversight. Forfis builds these controls into the pilot from day one: every classification is logged, every decision is auditable, and no candidate is screened out without human review. This is not a compliance checkbox added at the end; it is the architecture of the system.

    3. Use pgvector for Grounded Answers

    pgvector is a PostgreSQL extension that stores vector embeddings and performs similarity search. For a 501-2000 employee company, this means you can run your RAG pipeline on the same database as your transactional data, avoiding the cost and complexity of a dedicated vector database. The assistant embeds your job descriptions, screening criteria, and past hiring decisions into pgvector. When a new application arrives, the system retrieves the most relevant chunks and feeds them to an LLM, which generates a classification grounded in your data. This reduces hallucinations and keeps answers current as your criteria change.

    4. Integrate With Your Existing ATS via REST APIs

    The assistant connects to your ATS, HRIS, or recruitment platform via their REST APIs. Webhooks trigger the screening workflow when a new application arrives. The system extracts structured data from resumes, classifies candidates, and writes results back to your existing system. No replacement of your current tools is required. The architecture is deliberately model-agnostic: OpenAI or Anthropic APIs where quality matters, open-weight models on your own hardware where regulated data cannot leave the building. This means you can switch models without rebuilding the pipeline, and you can keep candidate data within your infrastructure if required.

    5. Ship a Measurable Result in 4 Weeks

    A 4-week timeline is realistic for a single-process pilot. Week 1: process audit and baseline measurement. Week 2: build the RAG pipeline and API integration. Week 3: test with real data and tune the model. Week 4: measure results, document findings, and hand over. This assumes your APIs are accessible and your data is in a usable format. The pilot ships with a report showing whether the automation meets the agreed thresholds on cycle time and error rate before you commit to rollout. This fixed-scope approach protects you from scope creep and ensures you have a measurable result before expanding to other workflows.

    6. Keep Humans in the Loop for High-Risk Decisions

    The assistant drafts a shortlist of candidates based on your job description and screening criteria. A recruiter reviews each draft, approves or rejects the classification, and the system logs the decision. This human-in-the-loop design ensures no candidate is screened out without human review, satisfying EU AI Act requirements for high-risk AI systems. The model classifies, the person decides. This is not a limitation; it is the correct architecture for a regulated environment. Every pilot ships with a measured before/after baseline on cycle time and error rate, so you know exactly what the automation achieved and where human judgment still adds value.

  • Cutting Contract Review Errors by 60% in a Two-Week B2B SaaS Pilot

    1. Baseline Error Rate Is the Real KPI

    The finance team at a 2,000+ employee B2B SaaS company in Vienna processes roughly 1,200 contracts per month. Each one passes through a manual review queue where an analyst extracts termination clauses, liability caps, and auto-renewal flags into the ERP. The baseline error rate sits at 5.2%: a missed auto-renewal date or a misread liability cap ends up in the system and surfaces three months later during a renewal dispute. A two-week pilot with a dedicated AI team replaced the manual extraction step with a LangGraph pipeline that parses PDFs, extracts 14 structured fields, and writes the result to a staging table via a custom REST API. The measured error rate dropped to 1.8% on the pilot’s 300-contract sample, and cycle time per contract fell from 11 minutes to 90 seconds of model time plus 4 minutes of human approval. The pilot did not touch the production ERP; it ran on a read-only copy of the contract repository and output to a sandbox workspace in the CRM.

    2. LangGraph Handles the Multi-Step Extraction

    The extraction pipeline runs on LangGraph, not a single LLM call. The graph has five nodes: PDF ingestion (PyMuPDF for text-layer PDFs, Tesseract OCR fallback for scanned documents), clause segmentation (a fine-tuned classifier that splits the document into 8–12 logical sections), field extraction (GPT-4o for high-accuracy fields like liability caps, Llama 3 70B on the client’s own GPU for fields containing personal data), confidence scoring, and output formatting. The REST API endpoint POST /v1/extract accepts a multipart PDF upload and returns a JSON object with 14 fields, each carrying a confidence score between 0 and 1. Fields below 0.90 route to a human reviewer in the existing helpdesk queue; fields at or above 0.90 auto-populate the staging table. Webhooks fire on completion so the finance team’s dashboard updates without polling. The entire pipeline runs on the client’s AWS eu-central-1 region, keeping data within Austria’s borders.

    3. Two Weeks Is Enough for a Measured Pilot

    The pilot ran for exactly 14 calendar days. Days 1–3: process audit. The AI team shadowed three finance analysts, logged every manual step, and identified the 14 fields that caused the most downstream errors. Days 4–6: data preparation. The team pulled 300 historical contracts from the repository, had two analysts independently annotate the 14 fields, and resolved disagreements to build a gold-standard test set. Days 7–10: pipeline build and tuning. The LangGraph workflow was assembled, the extraction prompt was iterated four times, and the confidence threshold was calibrated so that the false-negative rate (a wrong value auto-approved) stayed below 0.5%. Days 11–14: measurement. The pipeline ran on the 300-contract set, and the team compared field-level accuracy against the gold set, measured cycle time, and produced a before/after report. The report included a cost model: at 1,200 contracts per month, the pilot’s error reduction translated to an estimated EUR 18,400 in avoided dispute costs per quarter.

    4. Human-in-the-Loop Is Non-Negotiable

    The model does not replace the analyst; it removes the 11 minutes of copy-paste and field-mapping that precede the actual judgment call. The human-in-the-loop design is explicit: the model drafts the 14 extracted fields, the analyst reviews them in a purpose-built UI that highlights low-confidence fields in amber, and the analyst approves or corrects before the record writes to the ERP. For a B2B SaaS company, the highest-risk fields are termination notice periods and liability caps, because a wrong value here has direct financial consequences. The pilot’s measurement showed that 78% of fields required no human correction, 19% needed a single-field edit, and 3% required a full re-extraction. The analyst’s role shifted from data entry to exception handling, which freed roughly 6.5 hours per analyst per week. The dedicated AI team operated the pipeline during the pilot, monitored confidence drift, and tuned the prompt when a new contract template appeared in the sample.

    5. The Integration Is a Thin REST Layer

    The pilot’s REST API and webhook architecture was designed to plug into the client’s existing stack without replacing it. The extraction service exposes a stateless POST /v1/extract endpoint that the finance team’s internal tool calls via a simple HTTP request. On completion, a webhook POSTs the result to the client’s CRM (Salesforce) and ERP (SAP S/4HANA) through their respective API endpoints. No middleware, no new database, no replacement of the existing document management system. The client’s IT team reviewed the API contract in day 2 of the pilot and approved the integration scope. The model-agnostic design meant the team could swap GPT-4o for Llama 3 on the client’s GPU for any field that contained personal data, without changing the API contract or the downstream integration. This matters for a 2,000+ employee firm where IT governance requires that no new SaaS dependency is introduced for a pilot that may not scale.

    6. What the Pilot Does Not Cover

    The pilot’s 1.8% error rate is not the end state. The team’s rollout plan, presented in the final pilot report, targets a 0.9% error rate within 90 days of production deployment. The path: expand the gold-standard test set from 300 to 2,000 contracts, add a second extraction pass for fields with confidence between 0.80 and 0.90, and introduce a feedback loop where analyst corrections are logged and used to fine-tune the clause-segmentation classifier. The dedicated AI team continues to operate the pipeline in production, monitoring a dashboard that tracks field-level accuracy, confidence distribution, and cycle time per contract. The B2B SaaS firm’s finance director approved the rollout on the basis of the pilot’s measured numbers, not a projection. The two-week window was sufficient because the scope was narrow: one document type, 14 fields, one team, one measurement. Expanding to multi-party agreements or adding a second document type (e.g., purchase orders) would require a second pilot of similar duration.

  • AI Lead-Qualification Agent for Professional Services: A 4-Week LangGraph Pilot

    The Lead-Qualification Bottleneck in Large Professional Services Firms

    In a 2,000+ employee professional services firm in the USA, lead qualification is a bottleneck that compounds. Inbound inquiries arrive through web forms, email, and phone. A business development rep or account executive must read each one, cross-reference the prospect’s firmographics in the CRM, check whether the firm is already a client, assess budget and timeline, and then decide whether to route the lead to a senior partner or to marketing nurture. This process takes 4 to 8 hours per lead on average. With 200 to 400 inbound leads per month, that is 1,600 to 3,200 hours of senior-staff time consumed by triage that does not require a partner’s judgment. The error rate on manual qualification—misclassifying a prospect’s industry, missing a conflict of interest, or overlooking a budget signal—runs 12 to 18 percent, which means qualified leads sit in nurture for days while unqualified ones consume partner attention. The affected roles are business development managers, account executives, and in some firms, junior associates who are not yet billable. The systems involved are the CRM (Salesforce, HubSpot, or a custom platform), the marketing automation tool (Marketo, HubSpot Marketing, or Braze), and the helpdesk or ticketing system where inbound inquiries first land. The metric that matters is cycle time from inbound inquiry to qualified-lead handoff, and the current baseline is measured in hours, not minutes.

    Why Off-the-Shelf Chatbots and Rules-Based Triage Fall Short

    The first common approach is to add more business development headcount. This scales linearly: double the leads, double the triage time. It does not reduce the per-lead cycle time, and it increases the error rate because new hires are less familiar with the firm’s client base and conflict-of-interest rules. The second approach is to deploy a rules-based chatbot on the website. These bots follow a fixed decision tree: “What is your budget?” “What is your timeline?” They cannot handle ambiguous answers, cannot look up the prospect’s existing relationship with the firm in the CRM, and cannot escalate to a human when the conversation goes off-script. The third approach is to use a generic LLM wrapper—prompt an API with the lead’s text and ask it to classify. This works for simple cases but fails when the classification depends on data that is not in the prompt: the prospect’s existing CRM record, the firm’s service-line matrix, or the current capacity of the relevant practice group. Without retrieval-augmented generation grounded in the firm’s own data, the model hallucinates firmographic details and produces qualification scores that are not auditable. None of these approaches integrate with the existing CRM and marketing automation stack; they create a parallel system that the sales team must manually reconcile, adding friction rather than removing it.

    A LangGraph-Based Conversational Agent with Human-in-the-Loop Approval

    The alternative is a conversational agent built on LangChain and LangGraph, integrated through custom REST APIs and webhooks into the firm’s existing CRM, marketing automation, and helpdesk. LangGraph models the qualification workflow as a stateful graph: each node is a step (classify intent, retrieve the prospect’s CRM record, ask a follow-up question, score the response, draft a handoff summary), and edges define conditional transitions based on the prospect’s answers. The agent uses a model-agnostic architecture: OpenAI or Anthropic APIs for the conversational layer where response quality matters, and an open-weight model on the firm’s own hardware if any part of the data cannot leave the building due to client confidentiality agreements. The agent is human-in-the-loop by default: it drafts the qualification decision, a designated approver reviews it in a lightweight dashboard, and only after approval does the CRM record update and the webhook fire to the marketing automation tool. The pilot ships with a measured before/after baseline on cycle time and error rate, and the architecture is ISO 27001-aligned: all prompts and responses are logged, PII is encrypted, and access to the agent’s admin console is role-based. The delivery model is managed AI operations: the vendor operates the agent in production, monitors latency and error rates, and tunes prompts quarterly as the firm’s qualification criteria evolve.

    Four Concrete Steps to Start the Pilot

    Week 1 is the process audit. Map every inbound channel (web form, email, phone, referral), document the current triage steps, identify the CRM fields the agent will read and write, and define the qualification criteria as a structured rubric (industry, firm size, budget range, timeline, conflict-of-interest check). Confirm the ISO 27001 requirements: what data can be sent to an external API, what must stay on-premises, and what the audit log must capture. Week 2 is the build. Stand up the LangGraph agent, connect the custom REST APIs to the CRM and marketing automation tool, and implement the webhook that fires when a lead is marked qualified. Set up the human-in-the-loop approval queue with a 15-minute SLA. Week 3 is internal testing. Run 50 to 100 synthetic conversations covering edge cases: a prospect who is already a client, a prospect who asks for a specific partner, a prospect who gives an ambiguous budget answer. Measure the agent’s accuracy against the rubric and tune the prompts. Week 4 is the soft launch. Route 10 percent of live inbound leads through the agent, monitor the cycle time and error rate in real time, and document the before/after comparison. Full rollout to 100 percent of leads adds 2 to 4 weeks after the pilot, depending on the firm’s change-management process.