Tag: Lead Qualification

  • On-Premise AI Lead Qualification for a Swiss Professional Services Firm

    The Problem: 52-Hour Response Gaps and 6-Hour Reporting Cycles

    A 51-200 person professional services firm in Switzerland faces a specific operational bottleneck: inbound inquiries arrive across time zones and channels, but the sales team works 09:00-17:00 CET, Monday through Friday. A lead that lands at 22:00 on a Thursday waits 52 hours for a first substantive reply. In B2B professional services, that gap is not a minor inconvenience; it is a measurable conversion loss. The firm’s CRM holds the pipeline data, its Notion workspace holds the methodology documents, pricing sheets, and case studies, and its monthly reporting cycle consumes roughly 6 analyst-hours per month assembling numbers that already exist in the CRM.

    The problem is not a lack of data. It is a lack of a system that reads the data, classifies the inquiry, drafts a response, and files the report without a human touching each step. The firm does not need a new CRM or a new helpdesk. It needs an intelligent layer that sits on top of the tools it already runs, operates around the clock, and keeps every data point inside its own infrastructure because Swiss data-protection expectations and GDPR Article 32 make off-premise processing of client and prospect data a compliance risk the firm is not willing to take.

    Mechanism: On-Premise RAG, Open-Weight LLM, and the CRM Integration Layer

    The architecture has three components: a retrieval-augmented generation (RAG) pipeline, a conversational agent, and a reporting module. All three run on the client’s own hardware.

    The RAG pipeline ingests documents from the firm’s Notion workspace via the Notion API (version 2022-06-28), which exposes pages and blocks as JSON. Documents are chunked at heading boundaries, embedded with a sentence-transformer model (e.g., all-MiniLM-L6-v2, 384-dimensional vectors), and stored in a local Qdrant instance. At query time, the agent retrieves the top-5 chunks, builds a prompt with the retrieved context, and calls an open-weight LLM—Llama 3 70B or Mistral 8x7B—running on the firm’s GPU server. No document content or query text leaves the building.

    The conversational agent classifies each inbound inquiry into tiers: high-intent, mid-intent, low-intent. High-intent leads are routed to the CRM via its REST API with a structured summary. A human reviews every high-intent classification before the CRM record is created. The reporting module ingests CRM pipeline data and the firm’s Notion templates, drafts a structured monthly report with variance analysis, and queues it for human approval.

    The model-agnostic design means the firm can swap the LLM backend if a newer open-weight model outperforms the current one, without changing the RAG pipeline or the CRM integration.

    Trade-offs: Model Quality, Human Oversight, and Timeline

    The first trade-off is model quality versus data residency. A frontier API model (GPT-4o, Claude 3.5 Sonnet) would produce more nuanced lead classifications and better report narratives. But sending prospect names, firm details, and inquiry text to a third-party API violates the firm’s data-residency policy and complicates the GDPR Article 28 processor assessment. The open-weight model on-premise trades roughly 10-15% in classification accuracy for full data control. For a 51-200 person firm where the sales team reviews every high-intent lead anyway, that accuracy gap is acceptable.

    The second trade-off is the human-in-the-loop gate. Every high-intent classification requires a human approval before the CRM record is created. This adds roughly 90 seconds per lead and means the agent cannot fully automate the pipeline. But it eliminates the risk of a misqualified lead consuming a senior consultant’s time, and it satisfies the firm’s internal governance requirement that no AI output touches the sales pipeline without human sign-off.

    The third trade-off is the 4-week timeline. A full production rollout with monitoring, alerting, and a second channel would take 8-10 weeks. The 4-week pilot scopes to one workflow—lead qualification from inbound inquiries—and ships with a measured before/after baseline on cycle time and error rate. The firm accepts a narrower scope in exchange for a faster proof of value.

    Recommendation: Scope the Pilot to One Workflow, Measure the Delta

    The pilot targets lead qualification from inbound inquiries. The process audit in Week 1 maps the current workflow: inquiries arrive via email, web form, and phone, a sales associate manually classifies each one, drafts a first response, and logs the lead in the CRM. The baseline measurement captures cycle time (median 38 hours from inquiry to first response) and error rate (12% of leads misclassified in the prior quarter).

    Week 2 builds the RAG pipeline and connects the Notion API. Week 3 runs the agent in shadow mode against 200 historical inquiries, comparing its classifications to the human baseline. Week 4 adds the approval gate, connects the CRM write path, and measures the after-state. The target: reduce median first-response time to under 15 minutes for round-the-clock inquiries, and reduce misclassification rate to under 5%.

    The monthly reporting module ships in the same pilot. It ingests CRM pipeline data and the firm’s Notion reporting templates, drafts the monthly report, and queues it for analyst review. The target: reduce assembly time from 6 hours to 45 minutes of review and editing.

    The dedicated AI team operates as an embedded unit. The firm’s engineers and operations staff work alongside the team daily, not through a ticketing queue. This matters for a 4-week timeline: the team needs direct access to the Notion workspace, the CRM API credentials, and the firm’s GPU server, and it needs the operations staff available for the shadow-mode testing in Week 3.

  • Three Months to Cut Back-Office Errors in a German Fintech

    1. Start with a Process Audit, Not a Model

    A German fintech with 30 employees processes 400 payment-related documents per week. The back-office team spends 12 hours a week manually extracting data from invoices and payment confirmations, with a 4% error rate that triggers reconciliation delays. Forfis starts with a process audit that maps every manual touchpoint, then selects document extraction as the pilot workflow. The fixed-scope pilot runs for six weeks, shipping with a measured baseline: cycle time drops from 18 minutes per document to 4 minutes, and the error rate falls to 0.8%. The pilot’s success criteria are explicit and tied to the audit’s findings, not vague “efficiency gains.”

    2. Run the Pilot on Document Extraction

    The pilot targets one workflow: extracting line items, amounts, and reference numbers from payment statements and invoices. Forfis uses an open-weight model on the client’s own hardware because PCI DSS requires cardholder data to stay within a controlled environment. The model runs on a single GPU server in the client’s Frankfurt data center. The extraction pipeline feeds directly into the existing ERP via API, so no new data store is introduced. A human reviews every extracted record before it posts to the ledger, satisfying the human-in-the-loop requirement for anything touching money.

    3. Layer a Lead-Qualification Assistant on the CRM

    With the back-office pilot validated, the second phase adds a customer-facing AI assistant for lead qualification. The assistant pulls from the CRM and a Confluence knowledge base to draft first-response emails for inbound leads. It classifies each lead by intent, budget range, and product fit, then flags high-value prospects for the sales team. A rep approves every outbound message before it sends. The assistant reduces initial qualification time from 25 minutes to under 5 per lead, and the sales team reports a 15% lift in response rate within the first month of rollout.

    4. Keep the Stack Model-Agnostic and On-Premise

    The architecture is deliberately model-agnostic. OpenAI and Anthropic APIs handle non-sensitive tasks like drafting marketing copy or summarizing meeting notes. Open-weight models on the client’s hardware handle anything touching payment data, health records, or contracts. This split lets the fintech use frontier models where quality matters most while keeping regulated data on-premise. The integration layer plugs into the existing CRM, ERP, and helpdesk through their native APIs, so no system is replaced. For a 30-person team, this means no new vendor lock-in and no migration project.

    5. Scale Across Departments in the Third Month

    After the pilot, the rollout extends to two adjacent departments: the finance team adopts the document extraction pipeline for vendor invoices, and the support team uses the same RAG assistant for ticket triage. The key is that each new workflow reuses the same architecture, the same on-premise model, and the same human-approval gate. Forfis ships a measured before/after baseline for every workflow: cycle time, error rate, and cost per transaction. By month three, the back-office error rate has dropped from 4% to 0.8% across all automated workflows, and the team has freed up roughly 20 hours per week for higher-value work.

    6. Ship a Measured Baseline, Not a Promise

    The three-month timeline works because the scope is fixed and the success criteria are measurable. The process audit takes two weeks, the pilot runs six weeks, and the rollout occupies the final four weeks. For a 30-person fintech in Germany, this means no open-ended engagement and no surprise invoices. The human-in-the-loop design means the team never has to trust the model blindly: anything touching money, contracts, or health data gets a human sign-off. The result is a back office that runs on 0.8% error rates, a sales team that responds to leads in under five minutes, and an architecture that keeps PCI DSS-compliant data on the client’s own hardware.

  • 12-Step Checklist: RAG Pilot for Lead Qualification in German Insurance

    Pre-Pilot: Scope and Compliance Setup

    A 2-week fixed-scope pilot in a German insurance firm must produce a working RAG assistant on one workflow, a GDPR-compliant data-flow document, and a measured before/after baseline. The checklist below is operational: each item is a task a team can mark done or not done. It assumes the team uses LangChain and LangGraph, integrates with Slack or Microsoft Teams, and targets lead qualification to cut first-response time. The pilot is not a production deployment; it is a scoped experiment with a clear exit criterion. Work through the items in order. Skipping the audit or the baseline measurement invalidates the pilot’s value as a decision input for rollout.

    Build the RAG Pipeline on LangChain and LangGraph

    The RAG pipeline is the core of the pilot. Build it on LangChain for document chunking, embedding, and vector search, and on LangGraph for the stateful workflow that routes queries, handles multi-turn context, and triggers the human-approval gate. Keep the graph simple: one retrieval node, one generation node, one approval gate. Use a managed vector store in an EU region for the pilot. If the client’s data cannot leave the building, switch to an on-premises vector store and an open-weight model on the client’s GPU hardware. The RAG code is identical; only the embedding and inference endpoints change. Test the pipeline against 20 real lead queries before integrating with Slack or Teams.

    Integrate with Slack or Microsoft Teams

    The pilot must integrate with the channel the team already uses: Slack or Microsoft Teams. Build a bot that receives the lead query, calls the RAG pipeline, and returns the draft qualification score and suggested next step. The bot must include a human-approval gate: if the AI’s confidence drops below a threshold, or if the lead involves health-related data, the bot flags the query for a human agent. Log every human override. The integration must not replace the existing CRM or helpdesk; it plugs into them via their APIs. For a 2-week pilot, use a single OpenAI or Anthropic API endpoint for the LLM layer. Keep the model-agnostic layer thin: a single abstraction over the API call so switching providers later requires only a config change.

    Measure the Before/After Baseline

    Before the pilot starts, measure the baseline: cycle time from lead entry to qualified status, and error rate (misclassified leads) for the 2 weeks prior. Document the sample size, the definition of ‘error,’ and the measurement method. During the pilot, measure the same metrics for the 2 weeks of the pilot. The before/after comparison is the pilot’s primary deliverable. Without it, the client has no objective basis for the rollout decision. The baseline report must include: the number of leads processed, the average cycle time before and after, the error rate before and after, and the number of human overrides. This report is the exit criterion for the pilot.

    Define the Pilot Exit Criterion

    The pilot is a fixed-scope engagement: the vendor delivers a defined set of artifacts within the 2-week deadline. It is not a subscription or managed service. After the pilot, the client decides whether to proceed to rollout. The pilot includes a measured before/after comparison on cycle time and error rate, giving the client objective data to justify or reject the full deployment. The exit criterion is clear: if the pilot reduces cycle time by at least 30% and error rate by at least 20%, the client proceeds to rollout. If not, the pilot ends, and the client retains the baseline report and the RAG pipeline code. The vendor does not retain any client data after the pilot ends.

    Maintain the Checklist Over Time

    The checklist is a living document. After the pilot, review each item: mark what worked, what did not, and what needs adjustment. If the pilot proceeds to rollout, update the checklist to reflect the new scope: additional workflows, multi-language support, production monitoring. If the pilot ends, archive the checklist with the baseline report. Revisit the checklist before any new pilot: the GDPR landscape, the LLM provider landscape, and the integration landscape change. The checklist is not a one-time artifact; it is a tool for continuous improvement in AI-native operations. Keep it in the team’s project management tool, not in a static PDF.