Tag: Internal Knowledge Search

  • How an Austrian Medtech Firm Cut Compliance Reporting from 16 Hours to 4

    Background: A 30-Person Austrian Medtech Firm

    This case study is a composite built from patterns Forfis has observed across multiple engagements in the field. No named customer appears. The company, metrics, and timeline are representative of a recurring profile: a mid-size medtech firm in a Tier-1 European market running isolated AI pilots and looking to consolidate them into a managed workflow.

    The company is a 30-person Austrian medtech firm, roughly 18 months post-Series A, selling a Class IIa diagnostic device across DACH and Benelux. Its stack is a mix of a legacy CRM, a document management system for contracts, and a spreadsheet-driven compliance calendar. The compliance function is two people: a head of legal and compliance and a junior analyst. Monthly reporting under the EU Medical Device Regulation (MDR) and ISO 27001 requires them to extract obligations from 40+ active contracts, score each against operational status, flag deviations, and file a narrative summary with the quality management system. The cycle takes 14-18 hours per month, and the junior analyst is the single point of failure.

    Challenge: A 16-Hour Monthly Cycle and a Departing Analyst

    The pressure was operational, not strategic. The junior analyst was leaving in 90 days. The head of compliance had no bandwidth to absorb the full reporting cycle. The company was also preparing for an ISO 27001 surveillance audit in 5 months, which required documented, repeatable processes for every compliance activity. A manual, spreadsheet-driven cycle did not meet the audit’s evidence requirements.

    The specific need was to automate the monthly reporting cycle: extract clause-level obligations from contracts, score each against current operational status, flag deviations, and draft the narrative summary. The company had run two isolated AI pilots in the prior year — a ticket triage bot on their helpdesk and a document extraction tool for purchase orders — but neither touched the compliance function. The pilots were running, but they were not integrated, and the compliance team had no visibility into them. The challenge was not to build another isolated pilot but to create a managed, auditable workflow that the compliance team could own.

    Approach: Audit, Fixed-Scope Pilot, and a Dedicated Team

    Forfis ran a process audit in weeks 1-3. The audit mapped every step of the monthly reporting cycle, identified 5 automatable steps, and scored each by volume, error rate, and regulatory sensitivity. The pilot scope was fixed: clause extraction and obligation scoring for one product line, using the OpenAI API for text processing. The architecture was model-agnostic — the same pipeline could swap to an open-weight model on the client’s own hardware if a future contract contained data that could not leave the building. The integration was a custom REST API with webhooks, plugging into the existing CRM and document management system without replacing them.

    The delivery model was a dedicated AI team: three engineers and one product designer embedded with the compliance function for the full 6-month engagement. The team owned the model pipeline, the API, and the tuning loop. The compliance officer owned the approval step and the final report. Every pilot shipped with a measured before/after baseline on cycle time and error rate. The human-in-the-loop design meant the model drafted, the compliance officer approved, and every output was logged for the ISO 27001 audit trail.

    Outcome: 60% Cycle-Time Reduction and a Clean Audit

    The pilot ran for 6 weeks. The before baseline: 14-18 hours of manual work per month, with an error rate of 4-6% of obligations misclassified or missed. The after baseline at the end of the pilot: 4-6 hours of manual review per month, with an error rate of 1-2%. The cycle time dropped by roughly 60%. The compliance officer reported that the draft summaries were accurate enough to use as a starting point, cutting the drafting phase from 3 hours to 45 minutes.

    The go/no-go gate at the end of the pilot passed. Rollout extended the pipeline to all product lines and added the internal knowledge search layer, which indexed the contract corpus, regulatory guidance, CRM records, and past compliance reports. The knowledge search layer turned the monthly cycle into a continuous, queryable knowledge base. The team could answer ad-hoc questions like ‘What are our current MDR obligations for device X?’ in minutes instead of hours. The ISO 27001 surveillance audit, conducted in month 5, passed without findings on the reporting process. The dedicated team transitioned to a managed-operation model in month 7, handling model updates, prompt refinement, and edge-case triage on a monthly cadence.

    Lessons for Similar Teams

    • The audit is the product, not the pilot. The 3-week process audit produced a one-page decision matrix that the compliance team still uses 18 months later. The pilot was the validation, but the audit was the durable deliverable. Teams that skip the audit and jump straight to a pilot end up automating the wrong workflow.

    • Human-in-the-loop is not a compromise; it is the architecture. The model drafts, the human approves. This separation kept the ISO 27001 audit trail clean and ensured the model was never the final authority on a compliance determination. Teams that try to remove the human step for speed end up with an audit finding and a rework cycle.

    • Model-agnostic is a design constraint, not a marketing claim. The pipeline was built so the OpenAI API could be swapped for an open-weight model on the client’s own hardware without rewriting the integration. This mattered when a future contract contained data that could not leave the building. Teams that hard-code a single model API end up with a 3-month rework project when the data classification changes.

    • The knowledge search layer is where the ROI compounds. The monthly reporting cycle was the entry point, but the internal knowledge search layer is what the compliance team uses daily. The reporting cycle runs once a month; the knowledge search runs 20-30 times a week. Teams that stop at the reporting cycle miss the compounding value.

  • Fixed-Scope AI Pilot vs. Full Rollout: A Fintech’s 6-Month Decision

    What Is Being Compared: Fixed-Scope Pilot vs. Full-Scale Rollout

    The two options under comparison are a fixed-scope pilot and a full-scale rollout of AI automation across a 2,000+ employee fintech firm in the UK. The pilot targets one workflow — in this case, monthly reporting compilation and internal knowledge search over Notion and Confluence — with a 6-week delivery window, a measured before/after baseline on cycle time and error rate, and a go/no-go decision at the end. The full-scale rollout deploys AI process automation across multiple departments simultaneously: invoice processing, ticket triage for round-the-clock customer response, HR and recruiting workflow orchestration, and a retrieval-augmented assistant over the company’s documentation. Both options use the same underlying architecture: n8n for workflow orchestration, a model-agnostic AI layer (OpenAI or Anthropic APIs for non-regulated data, open-weight models on the client’s hardware for PCI DSS-sensitive data), and human-in-the-loop approval for anything touching money, contracts, or health data. The difference is scope, timeline, and risk exposure.

    Criteria for the Comparison

    The following criteria determine which option fits a fintech firm’s constraints. PCI DSS compliance is the hard gate: any workflow that touches cardholder data must run on-premises or in a PCI-compliant enclave, which rules out cloud-only model APIs for those specific flows. Cycle time reduction is measured in hours per report or per ticket, not in vague efficiency gains. Error rate is tracked as a percentage of transactions requiring manual correction. Integration depth counts the number of existing systems (CRM, ERP, helpdesk, Notion, Confluence) that the automation must connect to without replacing them. Vendor lock-in is assessed by whether the architecture can swap models or orchestration tools without rework. Timeline is the calendar duration from kickoff to managed operation. Cost is the total engagement fee plus ongoing managed operation, expressed in GBP. Scalability is the number of additional workflows or departments that can be added without rebuilding the core architecture.

    Comparison Table

    Criterion Fixed-Scope Pilot Full-Scale Rollout
    PCI DSS compliance One workflow isolated; open-weight model on-premises for cardholder data Multiple workflows; requires a PCI-compliant enclave for all payment-related flows
    Cycle time reduction Measured on one workflow (e.g., monthly reporting: 14 hrs → 2 hrs) Measured across 4-6 workflows; aggregate reduction depends on each workflow’s baseline
    Error rate Baseline established in week 1; target <2% by week 6 Baselines established per department; target <3% aggregate by month 4
    Integration depth 2-3 systems (Notion, Confluence, one CRM) 6-10 systems (CRM, ERP, helpdesk, Notion, Confluence, HRIS, payment gateway)
    Vendor lock-in Low; n8n workflows are portable; model can be swapped Moderate; more integrations increase switching cost, but n8n remains the orchestration layer
    Timeline 6 weeks to pilot completion; 2 weeks to decision 6 months to full managed operation across departments
    Cost (GBP) £18,000–£35,000 for the pilot £120,000–£250,000 for the full engagement plus £4,000–£8,000/month managed operation
    Scalability One workflow; scaling requires a new pilot per department Multi-department from day one; new workflows added to the existing n8n architecture

    When the Fixed-Scope Pilot Wins

    The fixed-scope pilot wins when the firm has not yet established a baseline for AI automation and needs to prove value before committing to a multi-department rollout. For a 2,000+ employee fintech in the UK, the pilot on monthly reporting and internal knowledge search over Notion and Confluence delivers a measurable result in 6 weeks: cycle time drops from 14 hours to 2 hours per report, and the error rate on data extraction falls from 8% to under 2%. The go/no-go decision is based on these numbers, not on a qualitative assessment. The pilot also validates the n8n orchestration layer and the human-in-the-loop approval gates without exposing the entire back office to change. If the pilot meets its targets, the firm has a proven template for the next workflow.

    The full-scale rollout wins when the firm has already completed a process audit, has identified 4-6 high-impact workflows, and has the IT capacity to manage parallel integrations. For a fintech with PCI DSS obligations, the rollout must include an on-premises open-weight model for any workflow that touches cardholder data, while non-regulated workflows (ticket triage, HR recruiting, knowledge search) can use OpenAI or Anthropic APIs. The 6-month timeline assumes that the process audit is complete, that the n8n environment is provisioned, and that each department has a named owner for the integration work. The rollout delivers aggregate cycle time reduction across the firm, but it requires a managed operation team from month 3 onward to handle model updates, integration drift, and new workflow requests.

    When the Full-Scale Rollout Wins

    The full-scale rollout is the right choice when the firm’s process audit has already identified multiple workflows with high impact and low integration complexity, and when the IT team can support parallel workstreams. For a 2,000+ employee fintech in the UK, this means the audit has scored invoice processing, ticket triage, HR and recruiting workflow orchestration, and internal knowledge search as the top four candidates. The rollout deploys all four within 6 months, with the PCI DSS-sensitive workflows (invoice processing, payment-related ticket triage) running on open-weight models on the client’s hardware, and the non-regulated workflows (HR recruiting, knowledge search) using OpenAI or Anthropic APIs. The n8n orchestration layer is shared across all workflows, so a change to one integration (e.g., a CRM API update) is applied once, not four times. The managed operation team, staffed from month 3, handles model retraining, integration monitoring, and new workflow requests. The cost is higher — £120,000 to £250,000 for the engagement plus £4,000 to £8,000 per month for managed operation — but the aggregate cycle time reduction across four workflows justifies the investment within 12 months for a firm of this size.

    Recommendation for the Scenario

    For a 2,000+ employee fintech in the UK with PCI DSS obligations, the recommendation is a fixed-scope pilot first, followed by a phased rollout. The pilot targets monthly reporting compilation and internal knowledge search over Notion and Confluence, with a 6-week delivery window and a measured baseline on cycle time and error rate. The pilot validates the n8n orchestration layer, the human-in-the-loop approval gates, and the model-agnostic architecture without exposing the payment processing workflows to change. If the pilot meets its targets — cycle time reduced from 14 hours to under 3 hours, error rate below 2% — the firm proceeds to a phased rollout over the remaining 4 months of the 6-month timeline. The rollout adds invoice processing, ticket triage for round-the-clock customer response, and HR and recruiting workflow orchestration, with PCI DSS-sensitive workflows running on open-weight models on the client’s hardware. The total engagement cost is £150,000 to £280,000, with managed operation at £5,000 to £8,000 per month from month 4 onward. This approach limits risk, delivers a measurable result in 6 weeks, and scales the architecture across departments without rebuilding it.

  • 2-Week AI Support Sprint for Swiss Professional Services Firms

    The Problem: Round-the-Clock Support Without Tripling Headcount

    Swiss professional services firms with 201-500 employees face a specific problem: customer support teams cannot provide round-the-clock coverage across German, French, Italian, and English without tripling headcount. The EU AI Act, which entered into force in August 2024, classifies customer support bots as limited-risk systems, requiring disclosure, human oversight, and documented evaluation metrics. A 2-week integration sprint addresses this by deploying an AI agent that handles first-response and ticket triage in Slack or Microsoft Teams, with a human-in-the-loop approval for anything touching money, health data, or contracts. The sprint delivers a measured baseline on cycle time and error rate, giving you concrete data before committing to full rollout. The architecture uses Anthropic Claude API for quality-critical tasks and open-weight models on local hardware for regulated data that cannot leave the building.

    Week 1: Process Audit and Prompt Engineering

    The sprint begins with a process audit that identifies which support workflows are worth automating. For a professional services firm, this typically includes ticket triage, first-response drafting, and internal knowledge search over CRM records and project documentation. The audit takes 2-3 days and produces a prioritized list of workflows ranked by volume, complexity, and compliance risk. The next phase is prompt engineering and API integration. The AI agent connects to your existing Slack or Microsoft Teams through their APIs, reads from your CRM and helpdesk, and drafts responses or classifies tickets. For multilingual coverage, the agent must be configured to detect the customer’s language and respond in German, French, Italian, or English as appropriate. Anthropic Claude supports 100+ languages, but you must ensure your internal documentation is available in each language for accurate retrieval-augmented responses.

    Week 2: Pilot Deployment and Baseline Measurement

    Week 2 focuses on the pilot and baseline measurement. The AI agent runs on one workflow, typically ticket triage or first-response drafting, while a human agent handles the same queue in parallel. The pilot measures cycle time (time from ticket creation to first response) and error rate (percentage of responses requiring human correction). For a professional services firm, a typical baseline shows a 40-60% reduction in cycle time and a 15-25% error rate for the AI agent, compared to 100% human handling. The human-in-the-loop approval ensures that anything touching money, health data, or contracts requires human sign-off before the response is sent. The pilot ships with a written report documenting the baseline metrics, the EU AI Act compliance checklist, and a recommendation for full rollout or scope adjustment. This fixed-scope structure protects you from open-ended costs and gives you concrete data to evaluate performance before committing to additional workflows.

    Compliance: EU AI Act and Swiss Data Protection

    The EU AI Act requires that customer support bots disclose they are AI, maintain human oversight for sensitive queries, and document the model’s training data and evaluation metrics. For Swiss firms, the Swiss Federal Act on Data Protection (revFADP) also applies to any personal data processed in the support flow. The architecture is deliberately model-agnostic: Anthropic Claude API handles quality-critical tasks where the model’s reasoning matters, while open-weight models on the client’s own hardware handle regulated data that cannot leave the building. This dual approach satisfies both quality requirements and data residency constraints. The AI agent plugs into existing CRMs, ERPs, helpdesks, and messaging through their APIs rather than replacing them, so your team continues working in the interfaces they already use. This integration approach minimizes disruption and training overhead, which is critical for a 2-week sprint.

    Pricing and Scope: What the 2-Week Sprint Delivers

    The 2-week sprint delivers a working pilot with measured baseline metrics, not a production-ready system. Full rollout across all support channels typically adds 4-6 weeks and includes additional workflows, multilingual coverage, and managed operation. The sprint cost for a 201-500 employee professional services firm ranges from EUR 15,000 to EUR 30,000 depending on integration complexity. This covers the process audit, prompt engineering, API integration, and pilot with baseline metrics. Ongoing managed operation and rollout to additional workflows are separate engagements with recurring costs based on API usage or hardware maintenance. The fixed-scope structure means you can evaluate performance before committing to full rollout. If the pilot does not meet your targets, you can adjust the scope or terminate the engagement without open-ended costs. This structure protects you from the common pitfall of AI projects that expand in scope and cost without delivering measurable results.

  • 10-Point Checklist: LLM Integration for HR and Recruiting in German Healthcare

    1. Audit and Baseline Measurement

    Before writing a single line of code, map the current state of HR and recruiting workflows. Identify which tasks involve PHI, which touch money or contracts, and which are purely administrative. This audit determines where human-in-the-loop approval is mandatory and where full automation is safe. Document baseline cycle time and error rate for each candidate workflow. This step prevents scope creep and ensures the pilot targets workflows with measurable ROI.

    • Audit all HR and recruiting workflows for PHI exposure and manual effort.
    • Measure baseline cycle time and error rate for each candidate workflow.
    • Identify human-in-the-loop approval points for PHI, money, or contract actions.
    • Document data sources in existing CRMs, ERPs, and helpdesks.
    • Define success metrics for the fixed-scope pilot before development begins.

    2. Fixed-Scope Pilot Definition

    Select one workflow for the fixed-scope pilot, typically internal knowledge search or document extraction. This workflow must have clear success metrics and a defined approval point. Avoid multi-workflow pilots; they dilute focus and complicate measurement. The pilot should ship with a measured before/after baseline on cycle time and error rate. A single, well-defined workflow allows you to validate the architecture and compliance controls before scaling.

    • Select one workflow for the fixed-scope pilot (e.g., internal knowledge search).
    • Define clear success metrics tied to cycle time and error rate.
    • Identify the human-in-the-loop approval point for PHI or contract actions.
    • Scope the pilot to avoid multi-workflow complexity.
    • Document the pilot’s success criteria before development begins.

    3. LangGraph Orchestration Setup

    Build the orchestration layer using LangChain and LangGraph. LangGraph handles stateful, multi-step workflows where nodes represent LLM calls, tool executions, or human approvals. Insert a mandatory human-in-the-loop node before any PHI is processed. This structure supports the fixed-scope pilot by isolating the workflow into discrete, testable states. LangGraph’s stateful design ensures that every step is auditable and reversible, which is critical for HIPAA compliance.

    • Implement LangGraph for stateful, multi-step workflow orchestration.
    • Insert human-in-the-loop nodes before any PHI processing.
    • Define state transitions for each workflow step.
    • Log every state change for auditability and compliance.
    • Test each node in isolation before integrating the full workflow.

    4. Model Selection and Deployment

    For regulated data that cannot leave the building, deploy open-weight models on the client’s own hardware. Use OpenAI or Anthropic APIs only for non-PHI tasks where quality matters and data residency is less critical. The architecture remains model-agnostic, allowing you to swap providers based on cost, latency, or compliance requirements. This approach ensures HIPAA compliance while maintaining flexibility in model selection.

    • Deploy open-weight models on-premises for PHI processing.
    • Use OpenAI/Anthropic APIs only for non-PHI tasks.
    • Configure model-agnostic architecture to swap providers easily.
    • Ensure data residency for all regulated data flows.
    • Document model selection criteria for compliance and cost.

    5. API and Webhook Integration

    Configure custom REST API endpoints and webhooks to connect the AI layer to existing HR systems, CRMs, and ERPs. Avoid replacing these systems; instead, plug into their APIs to retrieve data, trigger actions, and log outcomes. This approach preserves existing integrations and reduces migration risk. By integrating through APIs, you enable faster document turnaround without disrupting current operations.

    • Configure REST API endpoints for data retrieval and action triggers.
    • Set up webhooks for real-time event notifications.
    • Integrate with existing CRMs, ERPs, and helpdesks via their APIs.
    • Log all API calls for auditability and compliance.
    • Test integration points in a staging environment before production.

    6. HIPAA Compliance Controls

    Ensure all data flows are logged, access-controlled, and auditable to meet HIPAA Security Rule requirements. Implement role-based access control for PHI data. Encrypt data in transit and at rest. Document all access and modification events. These controls are non-negotiable for HIPAA compliance and must be in place before the pilot goes live.

    • Implement role-based access control for PHI data.
    • Encrypt data in transit and at rest using industry-standard protocols.
    • Log all access and modification events for auditability.
    • Document compliance controls for HIPAA Security Rule requirements.
    • Conduct a compliance review before the pilot goes live.

    7. Pilot Measurement and Iteration

    Measure the pilot’s performance against the baseline metrics defined in step 1. Compare cycle time and error rate before and after the pilot. If the pilot meets or exceeds targets, proceed to rollout; if not, iterate on the workflow design or model selection. This measurement ensures that the pilot delivers measurable value before scaling to additional departments or use cases.

    • Measure cycle time and error rate after the pilot.
    • Compare results against the baseline defined in step 1.
    • Document lessons learned from the pilot.
    • Iterate on workflow design if targets are not met.
    • Plan rollout based on pilot results and stakeholder feedback.
  • Open-Weight RAG vs. Cloud LLM APIs: Swiss Insurance Knowledge Search

    What Is Being Compared

    The two options under comparison are: (A) a retrieval-augmented knowledge assistant built on open-weight models (Llama 3 70B or Mistral 8x7B) deployed on the client’s own hardware, integrated into Microsoft Teams or Slack; and (B) the same RAG architecture but powered by OpenAI GPT-4o or Anthropic Claude 3.5 Sonnet via their public APIs. Both options serve the same use case: internal knowledge search over policy documents, claims procedures, and regulatory updates for a 501–2,000-person insurance or insurtech firm in Switzerland. The pilot scope is identical in both cases: one workflow, four weeks, a measured before/after baseline on cycle time and error rate, and a human-in-the-loop approval layer for compliance-sensitive queries. The difference is where the model runs and what that implies for latency, cost, data residency, and accuracy.

    Criteria for Judgment

    We judge the two options against six criteria that matter for a Swiss insurance firm operating under GDPR and FINMA supervision:

    • Data residency and GDPR compliance: whether personal data or special-category data (Article 9) can leave the client’s infrastructure.
    • Latency: end-to-end response time from query to answer, measured in milliseconds.
    • Accuracy on domain-specific retrieval: measured as top-k recall on a 200-query test set drawn from the client’s actual policy documents.
    • Cost at pilot scale: total cost of ownership for the 4-week pilot, including infrastructure, API calls, and integration work.
    • Vendor lock-in: how easily the client can swap models or providers after the pilot.
    • Operational overhead: who manages model updates, prompt tuning, and pipeline maintenance during the managed operations phase.

    Comparison Table

    Criterion Option A: Open-Weight On-Premise Option B: Cloud LLM API
    Data residency All data stays on client hardware; no external transmission Data transmitted to OpenAI or Anthropic servers (US/EU regions)
    GDPR Article 32 compliance Satisfied by default; no third-party processor Requires DPA and SCCs; Article 9 data requires additional safeguards
    Latency (p95) 180–350 ms (local inference, 8x A100 or equivalent) 400–900 ms (network round-trip + inference)
    Top-k recall (200-query test) 82–88% 91–95%
    Pilot cost (4 weeks) CHF 18,000–25,000 (hardware amortized + integration) CHF 8,000–12,000 (API calls + integration)
    Vendor lock-in Low; model weights are open, pipeline is portable Medium; prompt engineering and fine-tuning tied to provider
    Operational overhead Client manages hardware; Forfaq manages pipeline Forfaq manages pipeline; client manages API keys and billing

    Scenario-by-Scenario Verdict

    When Option A wins: The client’s knowledge base contains GDPR Article 9 special-category data (health-related policy terms, claims involving medical records) or Swiss data-residency requirements mandate that no data leaves the building. In this case, the 15–30% accuracy gap is acceptable because the queries are retrieval-heavy—finding the correct policy clause or regulatory citation—rather than complex multi-step reasoning. The 180–350 ms latency is well within the 2-second threshold for a back-office agent waiting for an answer in Teams. The 4-week pilot fits because the hardware is already provisioned or the client has existing GPU infrastructure.

    When Option B wins: The knowledge base is purely internal (policy terms, claims procedures, FINMA regulatory updates) with no personal data, and the client prioritizes accuracy over data residency. The 91–95% top-k recall matters when the assistant is used for compliance review, where a missed citation has regulatory consequences. The lower pilot cost (CHF 8,000–12,000 vs. CHF 18,000–25,000) makes it attractive for a first engagement. The 400–900 ms latency is acceptable for a back-office workflow where the agent is not on a live customer call.

    Recommendation

    For a 501–2,000-person Swiss insurance firm with one process already automated and a 4-week pilot timeline, Option A (open-weight on-premise) is the recommended choice if the knowledge base includes any GDPR Article 9 data or if Swiss data-residency policy prohibits external transmission. The accuracy gap is manageable for retrieval-heavy queries, and the data-residency advantage is non-negotiable for compliance. If the knowledge base is purely internal and the client’s primary goal is reducing error rate in compliance review, Option B (cloud API) is the better fit for the pilot, with a clear migration path to on-premise if the client later expands the assistant to handle personal data. In both cases, the human-in-the-loop approval layer is mandatory, and the managed operations agreement covers pipeline maintenance, prompt updates, and a 4-hour SLA for critical issues from week 5 onward.

  • LLM Integration vs. Round-the-Clock Response for E-commerce Support in Germany

    What Is Being Compared

    The two options under evaluation are distinct in scope and intent. Option A: LLM integration into existing systems embeds AI capabilities into the workflows a 51-200 person e-commerce company already runs. This includes an internal knowledge search over product catalogs, return policies, CRM records, and SOPs, plus a voice agent that handles inbound customer calls for order status, shipping updates, and return initiation. The integration layer uses n8n orchestration with custom REST API and webhook connections to the existing CRM, order management, and helpdesk. The model-agnostic architecture routes queries to OpenAI or Anthropic APIs for high-quality responses, or to open-weight models on the client’s own hardware when data sensitivity demands it. The pilot runs for 3 months with a measured before/after baseline on cycle time and error rate.

    Option B: Round-the-clock customer response is a narrower, channel-specific deployment. It focuses exclusively on the voice agent handling inbound calls 24/7, with the internal knowledge search serving as a supporting retrieval layer. The scope excludes broader system integration; the voice agent connects to the order management system via REST API for real-time order data, but does not extend to document extraction, invoice processing, or data entry automation. The human-in-the-loop approval layer routes any request involving refunds, cancellations, or disputes to a human agent. The pilot measures call handling time, first-contact resolution rate, and escalation rate.

    Evaluation Criteria

    The following criteria determine which option fits a 51-200 person e-commerce company in Germany running isolated pilots with a dedicated AI team and a 3-month timeline:

    • Cycle time reduction: measured in seconds for voice agent responses and minutes for knowledge search lookups, compared against the current human baseline.
    • Error rate: percentage of incorrect or incomplete responses in the pilot period, with a target below 5% for factual queries.
    • Integration depth: number of existing systems connected via REST API and webhooks, and the complexity of the n8n orchestration workflows.
    • Cost per interaction: API call costs for LLM inference, speech-to-text, and text-to-speech, amortized over the expected monthly interaction volume.
    • Staff time freed: hours per week per support agent redirected from routine tasks to complex escalations and retention work.
    • Vendor lock-in: degree of dependency on a single LLM provider, measured by the effort required to swap models without rewriting orchestration logic.
    • Scalability headroom: whether the n8n workflow architecture supports expansion from one use case to multiple channels within 6 months without a full rebuild.
    • Human-in-the-loop overhead: percentage of interactions requiring human approval, and the additional latency this adds to the customer experience.

    Side-by-Side Comparison

    Criterion Option A: LLM Integration Option B: Round-the-Clock Response
    Cycle time reduction 40-60% reduction in documentation lookup time; voice agent handles routine calls in under 90 seconds vs. 4-6 minutes for human agents Voice agent handles routine calls in under 90 seconds; no knowledge search component, so documentation lookup time remains unchanged
    Error rate Target below 5% for factual responses; RAG grounding reduces hallucination risk on policy and product queries Target below 5% for order status and shipping queries; no RAG layer, so responses rely on real-time API data only
    Integration depth 4-6 systems connected via REST API and webhooks: CRM, order management, helpdesk, product catalog, SOP repository, vector database 2-3 systems connected: order management, CRM, and speech-to-text/text-to-speech pipeline; no vector database or document indexing
    Cost per interaction EUR 0.03-0.08 per knowledge search query; EUR 0.15-0.40 per voice agent call (including STT, LLM, TTS) EUR 0.15-0.40 per voice agent call; no additional knowledge search cost
    Staff time freed 8-12 hours per agent per week across support and operations roles 6-10 hours per agent per week, concentrated on inbound call handling
    Vendor lock-in Low: n8n orchestration is model-agnostic; swapping between OpenAI, Anthropic, or open-weight models requires prompt adjustments, not workflow rewrites Moderate: voice agent pipeline is tied to specific STT and TTS providers; swapping requires re-testing the entire call flow
    Scalability headroom High: n8n workflows extend to additional channels (email, chat) and use cases (invoice processing, document extraction) within 6 months Low: adding knowledge search or document automation requires a separate integration project
    Human-in-the-loop overhead 15-25% of interactions require human approval (refunds, disputes, contract-related queries) 20-30% of calls require human escalation (refunds, cancellations, complex disputes)

    When Each Option Wins

    Option A wins when the company’s primary bottleneck is fragmented knowledge and repetitive documentation work. A 51-200 person e-commerce team in Germany typically maintains product catalogs, return policies, shipping documentation, and internal SOPs across 3-5 systems. The internal knowledge search consolidates these into a single retrieval layer, reducing lookup time from 5-10 minutes to under 30 seconds. The voice agent handles the inbound call volume that would otherwise tie up senior staff. The n8n orchestration layer connects to the CRM, order management, and helpdesk via REST API and webhooks, so the AI layer plugs into existing infrastructure rather than replacing it. For a company running isolated pilots, this broader integration scope justifies the 3-month timeline because the pilot delivers two measurable outcomes: reduced documentation lookup time and reduced call handling time.

    Option B wins when the company’s primary bottleneck is inbound call volume and the team wants a focused, low-risk pilot. The voice agent handles 60-70% of routine inbound calls (order status, shipping updates, return initiation) without requiring a vector database or document indexing pipeline. The integration scope is narrower: 2-3 systems connected via REST API, no RAG layer, no document extraction. The 3-month timeline is more comfortable because the build scope is smaller. The trade-off is that documentation lookup time remains unchanged, and the pilot does not demonstrate the company’s readiness for broader AI integration. For a team in the “Running Isolated Pilots” maturity stage, this focused approach reduces implementation risk and provides a clear before/after baseline on call handling metrics.

    Recommendation

    For a 51-200 person e-commerce company in Germany with a dedicated AI team, a 3-month timeline, and a need to free senior staff from routine work, Option A (LLM integration into existing systems) is the stronger fit. The reasoning is threefold. First, the “Need: Free Senior Staff from Routine Work” dimension implies that the bottleneck is not just call volume but also the time senior staff spend on documentation lookups, policy verification, and cross-system data retrieval. Option A addresses both bottlenecks; Option B addresses only the call volume. Second, the “AiMaturity: Running Isolated Pilots” stage benefits from a pilot that demonstrates the company’s ability to integrate AI across multiple systems, not just one channel. The n8n orchestration layer with 4-6 system connections provides a foundation for scaling to additional use cases (invoice processing, document extraction) within 6 months. Third, the model-agnostic architecture and human-in-the-loop approval layer reduce risk: the pilot ships with a measured before/after baseline on cycle time and error rate, and any output touching money or contracts requires human sign-off. The cost premium of Option A over Option B is approximately EUR 8,000-15,000 in additional development time for the knowledge search RAG pipeline and vector database setup, which is offset by the 8-12 hours per agent per week freed across the support and operations teams.

  • AI Automation Audit and Pilot for Monthly Reporting in a UK Healthcare Firm

    The Problem: Manual Reporting and Fragmented Knowledge in a 51-200 Person UK Healthcare Firm

    You run a 51-200 person healthcare or medtech firm in the UK. Your monthly reporting cycle — pulling data from intake forms, candidate tracking sheets, and operational logs, then assembling it into a board-ready summary — takes a dedicated person three to four days each month. There is no AI in production yet. Your stack is Google Workspace, a CRM, and a handful of spreadsheets. You need round-the-clock customer response on your public channels and an internal knowledge search that lets any team member pull answers from your own documents without asking a specific person. The problem is not a lack of data; it is that the data sits in unstructured documents, email threads, and manual entries, and no one has a systematic way to turn that into a scored, searchable, report-ready output. The fix is a fixed-scope, four-week engagement that starts with a process audit, moves to a pilot on one workflow, and ends with a measured baseline you can use to justify rollout.

    Prerequisites: What You Need Before the Audit Starts

    Before Forfis engineers touch your systems, you need the following in place:

    • Google Workspace admin access for the domain where your team operates. Forfis engineers need read access to Gmail, Drive, and Calendar to map document flows and email-based intake. You do not need to grant write access during the audit.
    • A named internal owner with authority to approve scope changes and sign off on the pilot. This person should be the one who currently owns the monthly reporting cycle, not a proxy.
    • Two weeks of historical data from your last reporting cycle: the raw intake documents, the intermediate spreadsheets, and the final report. Forfis uses this to build the baseline and train the predictive scoring model.
    • A list of the top 10 questions your team asks repeatedly that currently require a human to answer. This becomes the seed set for the RAG assistant.
    • A decision on the pilot workflow. Forfis recommends picking the one with the highest cycle time and the clearest before/after metric. For most firms at your size, that is the monthly reporting assembly step.

    Step 1: Run the AI Process Audit and Build the Roadmap

    Forfis engineers spend the first five business days mapping your current workflow. They sit with the person who runs the monthly report, watch them pull data from each source, and log every manual step. The output is a process map showing where documents enter the system, how they are classified, where they sit in queues, and how the final report is assembled. They also run a document inventory across your Google Drive and Gmail, tagging each file by type, frequency, and owner. By the end of day five, you have a one-page decision matrix ranking your workflows by cycle time, error rate, and automation feasibility. The audit does not write code. It produces a prioritized roadmap with a recommended pilot workflow and a projected cycle-time reduction. You review the matrix with your internal owner and confirm the pilot scope before moving to step two.

    Step 2: Build the Internal Knowledge Search Assistant on Google Workspace

    Forfis engineers connect to your Google Workspace via the Google Workspace API and pull the last two months of relevant documents, emails, and calendar events. They build a vector index using OpenAI’s text-embedding-3-small model, storing embeddings in a managed vector database (Qdrant or Pinecone, depending on your data volume). The index covers your policy documents, past reports, onboarding guides, and any internal wiki you maintain. The RAG assistant is exposed through a simple web interface and a Google Chat app so your team can ask questions in the channel they already use. The model behind the assistant is GPT-4o via the OpenAI API, configured with a system prompt that enforces citation of source documents and a refusal to answer questions outside the indexed corpus. You test the assistant with your top 10 seed questions and adjust the retrieval parameters (top-k, similarity threshold) until answers are accurate and cited.

    Step 3: Implement Predictive Scoring for Monthly Reporting

    Forfis engineers take the historical data from your last three reporting cycles and build a predictive scoring pipeline. Each incoming document or data point is scored on three dimensions: category (e.g., clinical intake, commercial inquiry, internal ops), urgency (based on keywords and sender patterns), and completeness (whether required fields are present). The model is GPT-4o-mini via the OpenAI API, chosen for cost efficiency at your volume. The scoring output is a JSON object with a confidence score per dimension. Anything below a 0.85 confidence threshold is routed to a human reviewer in a Google Sheets queue. The reviewer approves, corrects, or rejects the classification, and that correction feeds back into the model’s training set for the next cycle. You set the threshold in a single configuration file; Forfis engineers tune it during the pilot based on your tolerance for false positives versus false negatives.

    Step 4: Run the Four-Week Pilot and Measure the Baseline

    The pilot runs in shadow mode for the first two weeks. The AI pipeline processes every document and data point that would normally go through your manual workflow, but the output is not used for the actual report. Forfis engineers compare the AI output against what your team would have produced manually, logging every discrepancy. In week three, the pipeline goes live: the predictive scoring model classifies incoming items, the RAG assistant answers internal queries, and the human-in-the-loop queue handles low-confidence items. Your team continues to produce the monthly report as usual, but now the AI has already drafted the data summary and flagged anomalies. In week four, Forfis engineers measure the before/after baseline: cycle time from document receipt to report completion, and error rate (misclassified or missing data points). The pilot report includes both numbers side by side, a list of every discrepancy found in shadow mode, and a go/no-go recommendation for full rollout. You review the report with your internal owner and decide whether to proceed.

    Common Pitfalls and How to Detect Them

    The most common failure mode is scope creep during the audit. The audit is fixed-scope and two weeks long. If you ask Forfis engineers to add a new workflow mid-audit, the timeline slips. Detect this by reviewing the decision matrix at the end of day five and confirming the pilot scope in writing before moving to step two.

    • Stale vector index. If you add new documents to Google Drive after the index is built, the RAG assistant will not find them. Detect this by running a weekly re-index job and checking the index size in the vector database dashboard. If the document count has not increased in two weeks, the job is failing.

    • Overly aggressive confidence threshold. Setting the threshold too high (e.g., 0.95) routes most items to human review, negating the automation benefit. Detect this by monitoring the queue length in Google Sheets. If the queue exceeds 30 items per day, lower the threshold to 0.80 and re-measure.

    • No baseline data. If you cannot provide two weeks of historical data before the pilot starts, Forfis engineers cannot build the before/after comparison. Detect this in the prerequisites check. If you are missing data, delay the pilot start rather than proceeding without a baseline.

  • Dedicated AI Team vs Fractional Consultant for Medtech Monthly Reporting

    What Is Being Compared

    The firm is a 51-200 person UK healthcare and medtech company that has automated one back-office process and now faces two parallel needs: a customer-facing AI assistant for ticket triage and first-response, and an internal knowledge search layer over its own documentation and CRM records. The operational constraint is clear — scale these capabilities without adding headcount. The two options under evaluation are a dedicated AI team embedded for a 6-month engagement and a fractional consultant model where a single senior engineer works part-time across multiple clients. Both use LangChain and LangGraph as the orchestration layer, integrate with Google Workspace APIs, and ship with a human-in-the-loop approval gate for anything touching patient data or contractual obligations. The comparison below judges them against eight criteria that matter to a compliance-sensitive medtech operator in Tier-1 markets.

    Criteria for Judgment

    The eight criteria below reflect the specific constraints of a UK medtech firm at one-process-automated maturity:

    • Time-to-first-value: how many weeks until the agent handles a real workflow end-to-end.
    • Compliance documentation: whether the delivery model produces the audit trail MHRA and UK GDPR Article 22 expect.
    • Model-agnosticism: ability to swap OpenAI or Anthropic APIs for an open-weight model on client hardware if data residency rules tighten.
    • Integration depth: quality of the Google Workspace API layer (Drive, Gmail, Calendar) and CRM/ERP connectors.
    • Human-in-the-loop design: how the approval gate is architected, not just whether it exists.
    • Before/after measurement: whether the pilot ships with a quantified baseline on cycle time and error rate.
    • Knowledge-search recall: measured against a 200-query test set drawn from the firm’s own SOPs and regulatory correspondence.
    • Post-launch ownership: who monitors drift, handles model updates, and manages the eval suite after the 6-month window closes.

    Head-to-Head Comparison

    Criterion Dedicated AI Team Fractional Consultant
    Time-to-first-value 4-6 weeks to a working pilot on monthly reporting 8-12 weeks; consultant splits time across 3-4 clients
    Compliance documentation Full audit trail: prompt versions, model outputs, human-approval logs, eval results Partial; documentation depends on consultant’s personal practice
    Model-agnosticism Architecture designed for swap; open-weight Llama 3 70B on client hardware tested in week 3 Typically locked to one vendor API; swap requires re-architecture
    Google Workspace integration Native: Drive indexing, Gmail classification, Calendar-aware scheduling Basic: Drive read-only; Gmail integration often deferred
    Human-in-the-loop gate State-machine approval node in LangGraph; configurable per document type Simple if/else check; harder to extend to new document types
    Before/after baseline Measured at week 2 and week 12; cycle time and error rate tracked per workflow Often omitted or measured once at handover
    Knowledge-search recall 91-94% on 200-query test set after tuning 78-85% typical; tuning limited by consultant availability
    Post-launch ownership 3-month managed operation included; drift monitoring, eval suite maintenance Handover document; client owns all post-launch work

    When Each Option Wins

    The dedicated team wins when the firm needs the monthly reporting agent to feed a regulatory submission or board pack within the 6-month window. The state-machine approval node in LangGraph, combined with the measured before/after baseline, produces the documentation trail that a UK compliance lead can defend to an auditor. The fractional consultant model struggles here because the consultant’s time is split; the compliance documentation step, which takes 2-3 days of focused work, often slips to the end of the engagement or is delivered as a template rather than a filled-in record.

    For the customer-facing ticket triage agent, the dedicated team’s Google Workspace integration depth matters. The agent classifies incoming tickets by urgency and regulatory relevance, drafts a first response using the firm’s approved language, and escalates anything involving patient safety to a human. First-response time drops from 4 hours to under 15 minutes for routine queries. The fractional consultant can build this, but the integration with Gmail and Drive is typically read-only at handover, meaning the agent cannot draft responses into the firm’s existing workflow without additional work.

    For internal knowledge search, the dedicated team’s 91-94% recall on a 200-query test set, drawn from the firm’s own SOPs and regulatory correspondence, is the differentiator. The fractional consultant’s 78-85% recall is acceptable for casual lookups but insufficient when a compliance officer needs to find a specific regulatory decision from 18 months ago. The dedicated team’s tuning process, which includes iterating on chunking strategy and embedding model selection, is what closes that gap.

    Recommendation

    For a 51-200 person UK medtech firm at one-process-automated maturity, the dedicated AI team is the correct choice for a 6-month engagement covering monthly reporting, customer-facing ticket triage, and internal knowledge search. The reasons are specific: the compliance documentation requirement is non-negotiable in a healthcare context, the model-agnostic architecture protects the firm if data residency rules tighten, and the 3-month managed operation period after the 6-month build window means the firm is not left owning an eval suite and drift-monitoring pipeline it did not build. The fractional consultant model is appropriate for a firm that has already automated two or three processes and needs a single, well-scoped integration — not for a firm that is still at the one-process stage and needs the full audit-to-rollout lifecycle. The dedicated team’s EUR 18,000-25,000 per month cost over 6 months is comparable to the total cost of a fractional consultant at EUR 800-1,200 per day working 3-4 days per week, but the continuity of a named team and the built-in process-audit methodology make the dedicated model the lower-risk choice for a compliance-sensitive operator.

  • 8 Steps to Cut Back-Office Error Rates by 60-80% in 8 Weeks

    1. Measure the Baseline Before You Automate

    Before touching a single API, you need a documented baseline. For a 501-2000 employee B2B SaaS company, this means measuring the current cycle time and error rate for your target workflow—say, invoice processing or ticket triage. Pull 50-100 recent instances from your Zendesk or Intercom instance, timestamp each step, and log every error: misrouted tickets, duplicate invoices, missing fields. This baseline becomes your success metric. Without it, you can’t prove ROI or identify which model parameters need tuning. The audit also scores each workflow on volume, error cost, and automation feasibility, so you pick the one where a 20% error reduction saves the most money, not just the one with the highest volume.

    2. Scope the Pilot to One Workflow, Not a Platform

    The process audit identifies which workflows are worth automating, but the roadmap sequences them by ROI. For a B2B SaaS company, invoice processing often scores highest on error cost, while ticket triage scores highest on volume. The fixed-scope pilot then locks the deliverables: one workflow, one integration (Zendesk or Intercom), one success metric (error rate reduction), and an 8-week timeline. This bounded scope prevents scope creep and ensures you ship a measurable outcome. The pilot includes model configuration, API integration, human-in-the-loop approval workflow, and baseline measurement. You’re not building a platform—you’re proving that AI can cut error rates on one specific task before you scale.

    3. Use pgvector for Knowledge Search, Not a New Database

    For internal knowledge search, pgvector lets you store vector embeddings directly in your existing PostgreSQL database. You embed your documentation, CRM records, and support articles using OpenAI or Anthropic embedding models, then query them via similarity search. The advantage is operational simplicity: one database, one backup strategy, one access control layer. For a B2B SaaS company with 501-2000 employees, this means you don’t need a separate vector database like Pinecone or Weaviate. Latency for 100k vectors stays under 50ms on standard cloud PostgreSQL instances. The model-agnostic architecture means you can use commercial APIs for high-quality tasks and open-weight models on-premises when GDPR-regulated data cannot leave the building.

    4. Build Human-in-the-Loop Approval into the Workflow

    The model drafts or classifies, but a person approves anything that touches money, health data, or a contract. For a B2B SaaS company, this means the AI can auto-classify Zendesk tickets and draft first responses, but any output involving billing, customer data, or contractual terms requires manual approval before it’s sent. This hybrid approach gets you 80-90% of the automation benefit with 95%+ accuracy on high-stakes decisions. The approval workflow is built into the integration: the model flags items for review, a human approves or rejects, and the system logs every decision for audit. This keeps you GDPR-compliant under Article 22, which restricts automated decision-making with legal or similarly significant effects.

    5. Integrate with Zendesk or Intercom, Not a New Helpdesk

    The integration connects to Zendesk or Intercom’s API to pull ticket data, classify it using the AI model, and route it to the appropriate team or trigger a first-response draft. For document extraction, the system pulls invoices, contracts, or support articles from your existing systems, extracts key fields (PO numbers, dates, amounts), and validates them against your ERP or CRM. The model-agnostic architecture means you use OpenAI or Anthropic APIs where quality matters and open-weight models on the client’s own hardware where regulated data cannot leave the building. The integration plugs into your existing CRMs, ERPs, and helpdesks through their APIs, so you’re not replacing systems—just adding an AI layer on top. This keeps your existing workflows intact while cutting cycle time and error rates.

    6. Ship in 8 Weeks, Not 8 Months

    The 8-week timeline breaks down as: Week 1-2 (process audit and workflow selection), Week 3-4 (integration setup and model configuration), Week 5-6 (pilot deployment with human-in-the-loop approval), Week 7-8 (measurement, error rate analysis, and rollout planning). This assumes the client has API access to their Zendesk/Intercom instance and can provide 50-100 sample documents for training. Delays typically come from internal stakeholder alignment or data access permissions, not from the AI implementation itself. The pilot ships with a measured before/after baseline on cycle time and error rate, so you can prove ROI and identify which model parameters need tuning before you scale to additional workflows.

    7. Avoid the Five Most Common Pilot Failures

    The most common failure mode is skipping the baseline measurement. Without a documented before/after on cycle time and error rate, you can’t prove ROI or identify which model parameters need tuning. The second pitfall is automating a workflow with high decision complexity—like contract review—without a human-in-the-loop approval step. The third is underestimating integration work: Zendesk and Intercom APIs are well-documented, but mapping your ticket categories to model outputs and handling edge cases (malformed documents, missing fields) takes 2-3 weeks of engineering time that’s often overlooked in initial estimates. The fourth is choosing the wrong workflow: automate the one where a 20% error reduction saves the most money, not the one with the highest volume. The fifth is ignoring GDPR: if you’re processing EU customer data, you need a DPIA and audit logs, even for internal knowledge search.

  • AI Process Audit and 8-Week Integration Sprint for E-Commerce Support in the USA

    The Back-Office Bottleneck in a 2,000+ Employee E-Commerce Operation

    A 2,000+ employee e-commerce and retail company in the USA runs customer support across multiple channels: email, live chat, phone, and a self-service portal. The support team handles 15,000 to 25,000 tickets per month, with an average first-response time of 45 minutes and a misclassification rate of 12 percent. Back-office operations process 8,000 to 12,000 invoices monthly, with a data-entry error rate of 4 to 6 percent. Internal teams spend 3 to 5 hours per week searching through documentation, CRM records, and policy files to answer routine questions. The company has already automated one process, typically a document extraction workflow on the invoice pipeline, but the rest of the support and back-office stack still runs on manual triage, copy-paste data entry, and ad-hoc knowledge lookups. The pain is not a lack of tools. It is the absence of a measured baseline and a fixed-scope path from one automated process to a repeatable, auditable system that satisfies ISO 27001 controls.

    Why Off-the-Shelf Chatbots and In-House LLM Pipelines Fall Short

    Most companies at this stage reach for a generic chatbot platform or a point-solution RAG tool. The chatbot platform handles ticket routing but cannot access the company’s CRM, ERP, or internal documentation, so it deflects 60 to 70 percent of queries to a human agent without reducing cycle time. The RAG tool indexes a static document set but does not connect to live CRM records or helpdesk tickets, so the answers it returns are stale by the time a support agent reads them. A third common approach is to build a custom LLM pipeline in-house. This works for a single use case but requires a dedicated ML team, a GPU infrastructure budget of $15,000 to $40,000 per month, and 6 to 9 months of development before the first measurable result. None of these paths produce a fixed-scope pilot with a documented before/after baseline, which is the minimum evidence a CFO or compliance officer needs to approve a rollout. The failure mode is not technical. It is the absence of a delivery model that ties the build to a measurable outcome in 8 weeks or less.

    The Integration Sprint: Audit, Pilot, and Measured Baseline in 8 Weeks

    The integration sprint model starts with a process audit that maps every workflow in the support and back-office stack, measures cycle time and error rate on each, and ranks them by impact. The output is a fixed-scope pilot specification: one workflow, one integration, one measured outcome. For a company at the One Process Automated maturity stage, the next pilot is typically a conversational agent for customer support ticket triage or an internal knowledge search assistant built on retrieval-augmented generation over the company’s own documentation and CRM records. The architecture is model-agnostic: OpenAI or Anthropic APIs handle tasks where quality matters and data is non-sensitive, while open-weight models run on the client’s own hardware where regulated data cannot leave the building. The agent connects to the existing helpdesk, CRM, and ERP through their native REST APIs and webhooks. No system is replaced. The AI layer drafts, classifies, or retrieves; a human approves anything that touches money, health data, or a contract. The pilot ships with a documented before/after baseline on cycle time and error rate, which is the evidence the compliance team needs to map the new system to ISO 27001 Annex A controls.

    How to Start: Four Concrete Steps in the First 8 Weeks

    Week 1: run the process audit. Pull 90 days of ticket data from the helpdesk, 60 days of invoice data from the ERP, and a sample of internal knowledge queries from the support team. Measure cycle time, error rate, and volume on each workflow. Identify the two or three highest-impact candidates that can run in parallel without conflicting with the existing automation. Week 2: write the fixed-scope pilot specification. Define the target workflow, the integration points (which CRM fields, which helpdesk API endpoints, which document sources for the RAG index), the human-in-the-loop approval rules, and the before/after measurement plan. Week 3 to 5: build and integrate. Deploy the open-weight model on the client’s on-premise hardware for regulated data paths. Connect the agent to the helpdesk and CRM via REST API and webhooks. Build the RAG index over the company’s documentation and CRM records. Week 6 to 8: validate and measure. Run the agent in production with human approval on edge cases. Re-measure cycle time and error rate. Document the delta. Deliver the pilot report with the compliance mapping to ISO 27001 controls.