Tag: Free Senior Staff from Routine Work

  • RAG Candidate Screening for a 20-Person German Logistics Firm

    The Problem: Senior Staff Buried in Candidate Screening

    A 20-person logistics and supply chain company in Germany faces a recurring problem: senior operations managers spend 45 minutes per CV screening warehouse and fleet candidates, a task that scales linearly with applicant volume but adds no strategic value. The firm has no AI in production yet, no dedicated data team, and a hard constraint that personal data cannot leave German infrastructure due to GDPR. The need is not to replace HR but to free senior staff from routine work so they can focus on route optimization, supplier negotiations, and stakeholder management. The delivery model is a fixed-scope AI automation audit followed by a four-week pilot, with the goal of scaling operations without new hires. The use case is candidate screening, integrated with the firm’s existing Confluence documentation, and the AI stack is deliberately model-agnostic, using OpenAI’s API where quality matters and open-weight models on client hardware where regulated data cannot leave the building.

    How the RAG Assistant Works: Pipeline and Model Selection

    The system is a retrieval-augmented generation (RAG) assistant that ingests job descriptions, internal competency matrices, and past interview notes from Confluence via its REST API. The pipeline has three stages. First, a document parser extracts structured fields from CVs: name, contact, work history, certifications, and location. Second, a vector database (pgvector or Qdrant) stores embeddings of the job requirements and competency rubrics. Third, a language model scores each CV against the role’s requirements using a rubric defined by the hiring manager. The model is model-agnostic: OpenAI’s gpt-4o-mini handles non-personal tasks like formatting, while Llama 3 70B or Mistral 8x7B runs on the client’s own GPU server for any step touching personal data. The assistant drafts a shortlist with rationale and flags mismatches, such as a missing forklift certification for a warehouse role. A human reviewer approves or rejects each candidate before any communication goes out. The architecture is human-in-the-loop by default, and every pilot ships with a measured before/after baseline on cycle time and error rate.

    Trade-offs: Model Choice, Integration Depth, and Scope

    The architect faces three key trade-offs. First, model choice: OpenAI’s API offers higher quality for nuanced reasoning but requires a Standard Contractual Clause and data transfer to the US, which complicates GDPR compliance for personal data. Open-weight models on client hardware avoid this but require GPU infrastructure and tuning effort. For a 20-person firm, the cost of a single A100 GPU (roughly EUR 12,000 upfront or EUR 1,500/month via cloud) is justified if it eliminates the need for a data engineering hire. Second, integration depth: the assistant reads from Confluence via API but does not write back unless explicitly configured, preserving the existing governance model. This avoids the risk of the AI modifying source documents without human oversight. Third, scope: the pilot covers one hiring function, not the entire HR workflow. This keeps the four-week timeline realistic and the success criteria measurable. The trade-off is that the firm must decide which function to automate first, typically warehouse operations or fleet management, based on applicant volume and senior staff time spent.

    Recommendation: Audit, Pilot, and Rollout Path

    For a 20-person German logistics firm with no AI in production, the recommendation is a two-week audit followed by a two-week pilot on one hiring function. The audit maps the candidate screening workflow end-to-end, identifies which steps are rule-based versus judgment-based, and produces a prioritized automation roadmap. The pilot runs with a measured baseline: average time per CV, error rate on qualification decisions, and reviewer confidence. Success criteria are predefined: at least 40% reduction in screening time and no increase in false-positive rates. The architecture uses open-weight models on client hardware for any step touching personal data, with OpenAI’s API reserved for non-personal tasks. The assistant integrates with Confluence via its REST API, preserving existing access controls. The firm must provide candidates with information about the automated processing under GDPR Article 13 and 14, and the data processing agreement must specify that personal data is used for recruitment purposes only. The system does not make the final hiring decision; it reduces the time from 45 minutes per CV to under 5 minutes, freeing senior staff for strategic work.

  • AI Contract Review Agent vs. Back-Office Automation: UK Insurance Pilot

    What Is Being Compared

    The two options under comparison are distinct in scope and architecture. Option A is a purpose-built conversational AI agent for contract review, constructed on LangChain and LangGraph, that ingests insurance contracts via custom REST API and webhooks, extracts and classifies clauses, flags non-standard terms, and routes them for human approval. Option B is an extension of existing back-office automation, where the firm’s current invoice processing or document extraction pipeline is augmented with a lightweight classification layer to reduce manual review time without introducing a new conversational interface.

    Both options target the same business function: Finance and Accounting within an Insurance and Insurtech firm of 11-50 employees in the UK. Both must satisfy ISO 27001 controls and fit a 3-month fixed-scope pilot timeline. The difference lies in where the intelligence sits: Option A adds a reasoning layer that interprets contract language; Option B adds a pattern-matching layer that sorts documents faster.

    Criteria for Judgment

    The following criteria determine which option fits a 20-person UK insurance firm with ISO 27001 obligations:

    • Cycle time reduction: measured in hours per contract from receipt to approved status.
    • Error rate on clause classification: percentage of misclassified or missed non-standard clauses.
    • Integration effort: number of REST endpoints and webhook handlers required to connect to existing CRM, ERP, and document management systems.
    • Compliance overhead: additional controls needed to satisfy ISO 27001 Annex A requirements for data processing and audit logging.
    • Model dependency: whether the solution depends on a single commercial LLM API or can run on open-weight models on client hardware.
    • Scalability path: how the solution extends from one department to others without re-architecting.
    • Total cost of ownership over 12 months: including API fees, infrastructure, and internal staff time.
    • Change management burden: number of staff who must learn a new interface or workflow.

    Comparison Table

    Criterion Option A: Conversational Agent (LangGraph) Option B: Extended Back-Office Automation
    Cycle time reduction 40-50% for routine contracts; 20-30% for complex multi-party agreements 25-35% for document sorting; minimal for clause-level review
    Error rate on classification 1.5-3% with human-in-the-loop; 8-12% without 4-6% for document type; not applicable for clause semantics
    Integration effort 6-10 REST endpoints; 3-5 webhook handlers; 2-3 weeks build 2-4 REST endpoints; 1-2 webhook handlers; 1-2 weeks build
    ISO 27001 overhead Requires full audit trail of model prompts, outputs, and approvals; 2-3 additional Annex A controls Requires logging of classification decisions; 1 additional control
    Model dependency Can use OpenAI/Anthropic APIs or open-weight models on client hardware Typically rule-based or lightweight ML; no LLM dependency
    Scalability path Extends to new contract types by adding prompt templates and classification rules Extends to new document types by retraining classifier; limited semantic depth
    12-month TCO EUR 18,000-35,000 including API fees and infrastructure EUR 8,000-15,000 including maintenance
    Change management 3-5 staff learn new approval interface; 2-hour training 1-2 staff adjust sorting rules; 30-minute briefing

    Scenario-by-Scenario Verdict

    Option A wins when the firm’s bottleneck is clause-level interpretation. A 20-person insurance firm processing 150-300 contracts per month faces a specific problem: senior underwriters and finance staff spend 4-6 hours per contract reading, flagging, and summarizing terms. A conversational agent built on LangGraph can parse the contract, extract liability caps, renewal terms, and data processing clauses, and present a structured summary with confidence scores. The human reviewer then spends 30-45 minutes per contract instead of 4-6 hours. This directly addresses the need to free senior staff from routine work.

    Option B wins when the bottleneck is document volume, not complexity. If the firm’s problem is that 80% of incoming documents are routine renewals or endorsements that require minimal review, a classification layer that sorts them into “auto-approve” and “human review” queues reduces manual touchpoints without requiring semantic understanding. The integration is simpler, the compliance overhead is lower, and the 3-month timeline is easier to hit.

    Option A is the better fit for this scenario because the use case is explicitly contract review, not document sorting. The firm needs to understand what the contract says, not just what type of document it is.

    Recommendation

    For a UK insurance firm of 11-50 employees with ISO 27001 obligations, Option A — the conversational agent built on LangChain and LangGraph — is the recommended choice for the 3-month fixed-scope pilot. The reasoning is specific to the scenario dimensions:

    • The use case is contract review, which requires semantic understanding of clause language, not just document classification. Option B cannot flag a non-standard liability cap or an auto-renewal term buried in a 40-page policy.
    • The firm needs to free senior staff from routine work. A conversational agent that drafts summaries and flags exceptions reduces senior staff time by 40-50% on routine contracts, directly addressing this need.
    • ISO 27001 compliance is achievable with Option A if the architecture includes full audit logging of model prompts, outputs, and human approvals. The model-agnostic design allows the firm to use open-weight models on client hardware for sensitive policyholder data, keeping regulated data within the building.
    • The 3-month timeline is realistic: weeks 1-2 for process audit and baseline, weeks 3-6 for agent development and REST API integration, weeks 7-10 for human-in-the-loop testing, weeks 11-12 for documentation and handover.
    • Scaling across departments after the pilot is straightforward: the same LangGraph architecture extends to claims processing, underwriting, and customer service by adding new prompt templates and classification rules, without re-architecting the core agent.
  • AI Process Audit vs. Compliance-Safe Rollout for a German Logistics Firm

    What Is Being Compared

    The two options are distinct in scope and risk posture. Option A is an AI process audit and roadmap: a two-to-three-week engagement that maps the top 10 to 15 candidate workflows, scores them on volume, error rate, and integration complexity, and delivers a prioritized automation roadmap. The audit does not deploy any model. It produces a document: which workflows to automate, in what order, and with what expected cycle-time reduction. Option B is a compliance-safe AI rollout: a three-month engagement that includes the audit, a fixed-scope pilot on the highest-scoring workflow, rollout to the remaining high-impact workflows, and managed operations. The rollout ships a working AI layer integrated into the existing helpdesk and ERP, with a measured before/after baseline on cycle time and error rate. For a logistics and supply chain company in Germany with 201 to 500 employees, the decision hinges on whether the firm needs a plan or a working system by the end of the quarter.

    Criteria for Judgment

    The comparison rests on six criteria that matter to a mid-sized logistics operator in Germany. Time to first value: how many weeks until the firm sees a measurable reduction in manual work. Scope of deliverable: a document versus a running system. Integration depth: whether the option touches the existing SAP or Microsoft Dynamics ERP and helpdesk, or only recommends integration points. Risk exposure: the degree to which the option introduces a new AI layer into production before the firm has validated its accuracy. Cost structure: fixed-scope project fee versus ongoing managed operations retainer. Staff impact: whether the option frees senior staff from routine ticket triage and data cleanup within the three-month window, or defers that benefit to a later phase. Vendor lock-in: whether the architecture is model-agnostic and pluggable into existing systems, or tied to a single vendor’s platform. Compliance posture: whether the option includes a human-in-the-loop approval gate for any action that touches money, health data, or a contract, even when the firm’s own compliance requirements are minimal.

    Side-by-Side Comparison

    Criterion Option A: AI Process Audit Option B: Compliance-Safe Rollout
    Time to first value 2-3 weeks (roadmap delivered) 6-8 weeks (pilot live with baseline)
    Deliverable Prioritized workflow roadmap Working AI layer in helpdesk and ERP
    Integration depth Recommends integration points Live API integration with SAP/Dynamics
    Risk exposure None (no model deployed) Low (human-in-the-loop on all actions)
    Cost structure Fixed project fee, one-time Fixed pilot fee + monthly managed ops retainer
    Staff impact in 3 months None (plan only) Senior staff freed from routine triage by week 8
    Vendor lock-in None (document only) Model-agnostic; OpenAI/Anthropic or open-weight on client hardware
    Compliance posture N/A Human-in-the-loop; no regulated data leaves the building

    The table makes the trade-off explicit. Option A is cheaper and faster to deliver, but it produces no operational change within the three-month window. Option B costs more and takes longer to reach first value, but it delivers a working system that reduces cycle time and error rate by the end of the quarter.

    When Option A Wins

    Option A wins when the firm’s primary need is clarity, not speed. A logistics company with 201 to 500 employees that has not yet mapped its back-office workflows, or that is evaluating multiple automation vendors, benefits from a standalone audit. The roadmap becomes a procurement document: the firm can take the scored workflow list to three or four vendors and compare bids. The audit also suits a firm that expects to change its ERP or helpdesk within 12 months, because the roadmap can be re-scored against the new stack without re-running the full engagement. In this scenario, the three-month timeline is spent on the audit and internal decision-making, not on deployment.

    Option B wins when the firm’s primary need is operational relief within the quarter. A logistics operator whose senior staff are spending 15 to 20 hours per week on ticket triage, data enrichment, and cleanup for SAP or Microsoft Dynamics ERP records needs a working system, not a plan. The compliance-safe rollout ships a pilot on the highest-scoring workflow by week six, with a measured baseline showing cycle time and error rate before and after. By week twelve, the remaining high-impact workflows are live, and the managed operations retainer keeps the system running. The firm’s senior staff are freed from routine work within the three-month window, which is the stated need.

    When Option B Wins

    Option B wins when the firm’s primary need is operational relief within the quarter. A logistics operator whose senior staff are spending 15 to 20 hours per week on ticket triage, data enrichment, and cleanup for SAP or Microsoft Dynamics ERP records needs a working system, not a plan. The compliance-safe rollout ships a pilot on the highest-scoring workflow by week six, with a measured baseline showing cycle time and error rate before and after. By week twelve, the remaining high-impact workflows are live, and the managed operations retainer keeps the system running. The firm’s senior staff are freed from routine work within the three-month window, which is the stated need.

    Option A also wins when the firm’s compliance posture is genuinely minimal and the leadership team wants to defer the AI investment until the next budget cycle. The audit costs a fraction of the rollout, and the roadmap can be revisited in six months when the firm has more budget or a clearer strategic direction. However, this scenario is rare for a firm that has already identified ticket triage and data cleanup as the pain points. The stated need to free senior staff from routine work is an operational problem, not a strategic one, and it does not wait for the next budget cycle.

    Recommendation

    For a logistics and supply chain company in Germany with 201 to 500 employees, the stated need is to free senior staff from routine work within three months. The use case is ticket triage and routing, integrated with SAP or Microsoft Dynamics ERP, with data enrichment and cleanup as a secondary workflow. The firm has no specific compliance mandate beyond standard German data handling norms, and the delivery model is managed AI operations.

    Option B is the correct choice. The audit alone does not free any staff within the quarter. The rollout does. The compliance-safe rollout includes the audit as its first phase, so the firm gets the roadmap and the working system in the same engagement. The human-in-the-loop design means that no action touching money, health data, or a contract proceeds without a person’s approval, which addresses the risk concern even when the firm’s own compliance requirements are minimal. The model-agnostic architecture means the firm is not locked into a single vendor’s platform, and the integration with existing ERP and helpdesk APIs means no new infrastructure is required. The three-month timeline is sufficient: audit in weeks one to three, pilot in weeks four to eight, rollout in weeks nine to twelve, and managed operations from week twelve onward.

  • AI Process Audit vs. Compliance-Safe Rollout for B2B SaaS in the UAE

    What Is Being Compared

    Two distinct engagement models serve a 501-2000 employee B2B SaaS company in the UAE seeking to automate lead qualification and free senior staff from routine work. Option A: AI process audit and roadmap is a diagnostic engagement that maps existing workflows, measures baseline cycle time and error rate, and produces a prioritized automation roadmap. It does not deliver a working system; it delivers a plan. Option B: compliance-safe AI rollout is a fixed-scope pilot that implements one workflow end-to-end, with ISO 27001 controls, human-in-the-loop approval, and a measured before/after baseline. It delivers a working system on one workflow within a 4-week timeline. The two are not mutually exclusive: a typical engagement starts with Option A and proceeds to Option B, but they differ in scope, deliverables, and risk profile.

    Criteria for Comparison

    We judge both options against seven criteria that matter to a B2B SaaS company in the UAE with ISO 27001 obligations and a 4-week timeline:

    • Scope and deliverable: what the client receives at the end of the engagement.
    • Timeline fit: whether the engagement completes within 4 weeks.
    • Compliance readiness: how well the deliverable aligns with ISO 27001 controls.
    • Integration depth: how the deliverable connects to existing CRMs, ERPs, and Notion or Confluence.
    • Data handling: whether regulated data stays on client hardware or flows to external APIs.
    • Scalability: how easily the deliverable extends to additional departments.
    • Cost structure: fixed fee versus variable cost based on model usage.

    Comparison Table

    Criterion Option A: AI Process Audit and Roadmap Option B: Compliance-Safe AI Rollout
    Scope and deliverable Prioritized roadmap with 3-5 candidate workflows, baseline metrics, and pilot recommendation Working pilot on one workflow with measured before/after baseline on cycle time and error rate
    Timeline fit 2-3 weeks for audit and roadmap 4 weeks for pilot delivery, including ISO 27001 documentation and handover
    Compliance readiness Identifies compliance gaps and recommends controls; does not implement them Implements ISO 27001 controls: data classification, audit logging, human-in-the-loop approval
    Integration depth Maps existing APIs and identifies integration points Connects to CRM, ERP, helpdesk, and Notion or Confluence via their APIs
    Data handling Classifies data types and recommends routing (open-weight vs. API) Routes regulated data to open-weight models on client hardware; non-regulated data to OpenAI or Anthropic APIs
    Scalability Roadmap defines sequence for scaling across departments Pilot architecture reuses for adjacent departments, reducing integration cost
    Cost structure Fixed fee for audit and roadmap Fixed fee for pilot; variable cost for model usage during managed operation

    Scenario-by-Scenario Verdict

    Option A wins when the company has not yet identified which workflows to automate. A 501-2000 employee B2B SaaS company in the UAE may have 15-20 candidate workflows across marketing, sales, and operations. The audit narrows this to 3-5 high-impact workflows, such as lead qualification with data enrichment, document extraction from inbound forms, and ticket triage. The roadmap sequences these by ROI, ensuring the 4-week pilot targets the workflow with the highest measurable impact. Without this diagnostic step, the pilot risks automating a low-impact workflow and failing to demonstrate value.

    Option B wins when the company already knows which workflow to automate and needs a working system within 4 weeks. For a B2B SaaS company with ISO 27001 obligations, the rollout implements the compliance controls that Option A only recommends. The pilot ships with a measured before/after baseline on cycle time and error rate, providing the evidence needed to justify scaling to additional departments. The human-in-the-loop model ensures senior staff retain approval authority over outputs touching contracts or financial data.

    Recommendation

    For a 501-2000 employee B2B SaaS company in the UAE with ISO 27001 obligations and a 4-week timeline, the recommendation is to combine both options in sequence. Week 1 delivers the process audit and roadmap, identifying lead qualification with data enrichment as the highest-impact workflow. Weeks 2-4 deliver the compliance-safe AI rollout on that workflow, with pgvector embeddings search over Notion or Confluence documentation, model-agnostic routing (OpenAI or Anthropic APIs for non-regulated data, open-weight models on client hardware for regulated data), and human-in-the-loop approval for any output touching contracts or financial data. The dedicated AI team manages the full cycle, freeing senior staff from routine work while maintaining ISO 27001 compliance. This sequence ensures the pilot targets the right workflow and delivers a working system with measurable baselines within the 4-week constraint.

  • 12-Point Checklist: AI Lead-Qualification Pilot for a 20-Person UK Fintech Firm

    1. Verify the pilot scope is locked to one workflow

    Before any code is written, confirm the scope is locked to one workflow. For a 20-person fintech firm, that means the pilot covers lead qualification only — not invoice processing, not document extraction, not voice. The audit deliverable should name the specific CRM fields the agent will read and write, the webhook endpoints it will call, and the exact lead-qualification criteria the sales team already uses. A fixed scope prevents the pilot from drifting into a multi-week integration project that buries the team in configuration work instead of measuring cycle-time savings.

    2. Document the PCI DSS data-flow and risk assessment

    Run a formal risk assessment under PCI DSS Requirement 12.10 before the agent touches any production data. Document the data flow from first contact to qualified-lead status, confirm that cardholder data never enters the LLM prompt, and obtain a signed attestation from OpenAI that they do not retain training data. This documentation pack is a deliverable, not an afterthought. Without it, the pilot cannot pass internal governance review, and the 4-week timeline slips.

    3. Configure the CRM and webhook integration points

    Map every integration point before the pilot starts. The agent reads lead records from the CRM via its REST API, writes qualification scores back to the same CRM, and triggers webhooks to the helpdesk when a lead is flagged for human follow-up. Each endpoint needs an API key, a rate-limit budget, and a fallback path for when the CRM is down. For a 20-person firm, this typically means 3-5 endpoints, not 30.

    4. Define the human-in-the-loop approval threshold

    Set the human-in-the-loop threshold before the first test. The agent drafts the qualification response and classifies the lead, but a person approves any action that touches a contract, payment, or regulated data. For lead qualification, this means the agent can mark a lead as “qualified” or “unqualified” but cannot send a payment link or modify a contract clause. The approval step is logged with a timestamp and user ID, which feeds the error-rate baseline.

    5. Measure the before-state baseline on cycle time and error rate

    Capture the baseline before the agent goes live. Track time from first contact to qualified-lead status and the percentage of misclassified leads over a 2-week window using the existing manual process. These two numbers — cycle time and error rate — are the only metrics that matter for the pilot report. Everything else is noise. For a 20-person firm, a 2-week baseline is sufficient to establish a statistically meaningful before-state.

    6. Implement the cardholder-data filter and test it

    Build a pre-processing filter that strips or masks any field containing cardholder data, PAN, or CVV before the prompt is sent to the OpenAI API. Test the filter with synthetic data that includes edge cases: partial PANs, CVVs embedded in free-text notes, and card numbers in email subject lines. The filter must reject or flag any input that fails the mask, and the rejection log must be retained for the PCI DSS audit trail.

    7. Write and version the prompt template for lead qualification

    Write the prompt template that the agent uses to classify leads and draft responses. The template should include the firm’s specific qualification criteria, the tone of voice the sales team expects, and a clear instruction to reject any input that contains cardholder data. Version the prompt in a repository, not in a config file. Each change to the prompt should be logged with a reason, because prompt drift is the most common cause of error-rate spikes in the first two weeks of operation.

  • UAE Payments Firm Cuts Ticket Cycle Time 38% with a Claude-Based Triage Agent

    Background: A 2,400-Person Payments Firm in the UAE

    This case study is a composite drawn from patterns Forfis has observed across multiple engagements in Tier-1 markets. No named customer is represented. The details below reflect a recurring profile: a mid-to-large fintech or payments company in the UAE or Gulf region, operating under GDPR-equivalent data-protection rules, with a helpdesk that has outgrown manual triage.

    The company in this scenario is a payments processor with roughly 2,400 employees, a mix of engineering, compliance, and customer-operations staff. Its product stack includes a core payment engine, a merchant portal, and a customer-facing helpdesk running on a commercial platform. The helpdesk handles 18,000 to 22,000 tickets per month, the majority of which are routine: failed-payment inquiries, settlement-delay questions, and document-request follow-ups. Senior operations staff spend an estimated 35 to 45 percent of their week reading, categorizing, and routing these tickets before any substantive work begins.

    Challenge: Senior Staff Buried Under Routine Triage

    The operations director set a clear constraint: senior staff were being consumed by work that did not require their judgment. A payment-failure ticket that follows the standard runbook in Confluence should not be read by a team lead with eight years of settlement experience. The pressure was not just efficiency; it was retention. Three senior operations managers had left in the preceding year, citing repetitive triage as a primary factor.

    Compliance added a second constraint. The firm processes customer data subject to the UAE Data Protection Law (Federal Decree-Law No. 45 of 2021), which aligns closely with GDPR Articles 5, 28, and 30. Any AI system touching ticket content had to demonstrate data minimization, processor accountability, and a documented right-to-erasure path. The firm had already run two isolated pilots on document extraction for onboarding, but those pilots had not produced a measured baseline and had not moved to production. The operations team was skeptical of a third pilot unless the scope was narrow, the timeline was fixed, and the success criteria were written into the contract before a single line of code was written.

    Approach: A Fixed-Scope Pilot on One Workflow

    Forfis scoped the engagement as a fixed-scope, three-month pilot on a single workflow: ticket triage and routing for the payment-failure and settlement-delay categories. The architecture used the Anthropic Claude API for classification and summarization, with the model called from a lightweight service that read ticket content from the helpdesk’s REST API and wrote routing decisions back. The knowledge base lived in Confluence, queried through its search API to pull the relevant runbook for each ticket category.

    The delivery model was managed AI operations from day one. Forfis handled the technical planning, the prompt engineering, the evaluation harness, and the integration work. The client’s operations team provided the labeled sample set (400 historical tickets with correct routing decisions) and the Confluence content owners. The human-in-the-loop boundary was explicit: the agent classified and routed, but any ticket flagged as involving a refund, a contract amendment, or a regulatory report was suppressed from auto-routing and escalated to a senior reviewer. Every model call was logged with a retention window matching the firm’s records-management policy, satisfying the processor-accountability requirement under the UAE law and GDPR Article 30.

    Outcome: Measured Cycle-Time Reduction and Error-Rate Drop

    The pilot ran for twelve weeks. The first two weeks were the process audit: Forfis mapped the top ten ticket intents, measured the current median cycle time (4.2 hours from ticket creation to first substantive response) and the current misrouting rate (11.3 percent on a 300-ticket sample). Weeks three through six built the triage agent and the evaluation harness. Weeks seven through twelve ran shadow mode: the agent drafted a routing decision, a human approved or overrode it, and the override was logged.

    By week twelve, the agent’s classification accuracy on a held-out set of 200 tickets was 94.1 percent. The median cycle time for the two target categories dropped to 2.6 hours, a 38 percent reduction. The misrouting rate fell to 3.8 percent. Three senior operations managers reported spending roughly 12 to 15 hours per week less on initial triage, which they redirected to escalation handling and vendor-management work. The client extended the engagement to a managed-operations contract covering model monitoring, Confluence content review, and incident response at a fixed monthly fee. The pilot did not expand to fraud detection or chargeback handling; those remain separate engagements with their own baselines.

    Lessons for Teams Running Isolated Pilots in Regulated Sectors

    Five lessons from this engagement generalize to similar teams in regulated, high-volume operations:

    • Scope the pilot to one workflow, not a category. “Ticket triage” is too broad. “Triage and routing for payment-failure and settlement-delay tickets” is a contract. The narrower the scope, the more defensible the baseline and the faster the rollout decision.

    • Write the success criteria before the audit. The 94 percent accuracy threshold and the 30 percent cycle-time reduction were in the statement of work before Forfis touched the helpdesk API. Without that, the pilot becomes a demo, not a decision.

    • Keep the knowledge base in the tool the team already uses. Confluence was the source of truth for runbooks. Pulling from it via API meant the content owners did not need to learn a new system, and updates propagated without a retraining step.

    • Log every model call from day one. The compliance team asked for the audit trail in week four, not week twelve. Having it from week one turned a potential blocker into a non-issue.

    • Do not let the pilot absorb adjacent workflows. The operations team wanted fraud triage in week five. Holding the line kept the timeline realistic and the error-rate target achievable.

  • Document Extraction Pilot for E-Commerce Operations in Austria

    The Operational Bottleneck: Manual Order and Shipment Data Entry

    E-commerce and retail operations teams in Austria face a persistent bottleneck: order and shipment status updates from suppliers arrive in inconsistent formats—PDFs, scanned images, email attachments, and portal exports. Manual extraction and data entry into SAP or Microsoft Dynamics consumes 30-45 minutes per batch, with error rates averaging 2-4% that cascade into delayed customer notifications and reconciliation headaches.

    A fixed-scope pilot addresses this by automating one specific workflow within a four-week window. The engagement starts with a process audit that maps your current document flow, measures baseline cycle time and error rate, and identifies the highest-ROI extraction targets. From there, the team builds a document extraction pipeline using the OpenAI API for its strong performance on varied layouts, integrates it with your existing ERP via native APIs, and validates results against your baseline metrics.

    The deliverable is not a new system but a faster, more accurate version of the workflow you already run. Senior operations staff move from data entry to exception handling and supplier relationship management, while the AI layer handles the repetitive extraction and mapping work.

    Four-Week Pilot Structure: From Audit to Validated Pipeline

    The four-week timeline follows a structured sequence. Week one covers the process audit: the team reviews 50-100 sample documents from your supplier base, maps data fields to your ERP schema, and establishes the baseline metrics—current cycle time per batch, error rate, and staff hours consumed. This phase also confirms compliance requirements under the EU AI Act, including transparency logging and human oversight protocols for data that affects financial records.

    Weeks two and three handle model configuration and integration. The OpenAI API is tuned for your specific document types, with prompt engineering and post-processing rules to handle edge cases like merged invoices or multi-page shipments. The extraction pipeline connects to SAP or Microsoft Dynamics through their standard APIs, writing validated data directly to the relevant tables. Human-in-the-loop review queues are configured so that low-confidence extractions route to staff for approval before ERP sync.

    Week four focuses on validation and handover. The team processes a full week’s worth of live documents, compares results against the baseline, and documents the error rate, cycle time improvement, and any remaining edge cases. The handover package includes runbooks, model version records, and escalation procedures for ongoing managed operation.

    Model-Agnostic Architecture: OpenAI API and Open-Weight Options

    The architecture is deliberately model-agnostic, but the OpenAI API serves as the default for quality-critical extraction tasks. Its strength lies in handling varied document formats—scanned PDFs with mixed layouts, email attachments with inconsistent headers, and portal exports with variable column structures—without requiring custom OCR preprocessing for each format.

    For regulated data that cannot leave the building, the same pipeline runs on open-weight models deployed on your own hardware. This configuration maintains the same integration points and human-in-the-loop workflows while ensuring data sovereignty. The trade-off is higher initial setup effort and potentially lower accuracy on edge cases, which the human review queue compensates for.

    The pipeline plugs into your existing SAP or Microsoft Dynamics ERP through their standard APIs rather than replacing them. Extracted data maps to your existing data structures: order numbers to sales order tables, shipment dates to delivery schedule lines, status codes to your internal workflow states. No ERP migration or reconfiguration is required. The AI layer sits alongside your current systems, handling the extraction and mapping work while your ERP continues to manage the downstream business logic.

    EU AI Act Compliance: Transparency and Human Oversight

    The EU AI Act classifies document extraction systems as limited-risk AI, requiring transparency about AI involvement and human oversight for decisions that affect financial records or customer commitments. For e-commerce operations in Austria, this means the system must log its actions, maintain records of model versions and training data, and allow human review before extracted data syncs to the ERP.

    The pilot ships with compliance documentation built in: action logs showing which documents were processed, confidence scores for each extraction, and a review trail for any human approvals. Model version records track which API version or open-weight model was used for each batch, supporting audit requirements. The human-in-the-loop workflow ensures that anything touching money, health data, or contracts requires explicit staff approval before ERP sync.

    For a 501-2000 employee company, this compliance layer adds minimal overhead to the four-week timeline. The documentation and logging are configured during the integration phase, and the review queue is part of the standard human-in-the-loop setup. The result is a system that meets EU AI Act requirements without requiring a separate compliance project or legal review cycle.

    Measuring Success: Cycle Time, Error Rate, and Staff Hours

    The pilot’s success is measured against the baseline established in week one. Typical targets for order and shipment status extraction include reducing cycle time from 30-45 minutes per batch to under 10 minutes, cutting error rates from 2-4% to under 0.5%, and freeing 60-80% of the staff hours previously consumed by manual data entry.

    The before/after comparison uses the same document samples processed through both the manual and AI-assisted workflows. Cycle time measures the elapsed time from document receipt to ERP sync. Error rate counts the number of fields requiring correction after initial extraction, divided by total fields processed. Staff hours are tracked through time-stamped review queues, showing how much time staff spend on exception handling versus routine data entry.

    The handover package includes a validation report with these metrics, a runbook for daily operations, and escalation procedures for edge cases. The managed operation phase continues with monthly performance reviews, model updates as supplier document formats change, and support for new document types as your supplier base evolves. The goal is not a one-time automation but a continuously improving AI-native operations layer that scales with your business.

  • UK B2B SaaS Firm Cuts Invoice Cycle Time 74% With On-Premise AI Pilot

    Background: A 30-Person B2B SaaS Firm in Manchester

    This case study is a composite drawn from patterns observed across multiple engagements. No named customer is represented. The firm described here is a 30-person B2B SaaS company based in Manchester, selling a project-management platform to mid-market clients across the UK and Ireland. The operations team of four handles supplier invoices, delivery notes, and credit notes for a mix of cloud hosting, office supplies, and professional services vendors. The existing stack is a standard ERP (Xero for accounting, a lightweight project-management tool for internal tracking) and Slack as the primary communication channel. No AI system is in production anywhere in the company. The trigger for change is not a technology initiative but a headcount constraint: the operations lead has been absorbing invoice processing work that was previously split across two part-time staff, and the founder has set a deadline to reduce the manual workload before the next hiring cycle in Q3.

    Challenge: Four-Day Cycle Time and a GDPR Gap

    The operations lead processes roughly 180 supplier invoices per month, each requiring manual data entry into Xero: vendor name, line items, tax codes, and total amount. The median cycle time from invoice receipt to payment approval is four business days, with a long tail of invoices taking nine to twelve days when the operations lead is pulled into client escalations. The error rate on manual data entry is 18 percent, measured over a two-week sample in the audit phase. Each error triggers a correction cycle that adds 20 to 35 minutes of senior staff time. The compliance pressure is GDPR: the invoices contain personal data (vendor contact names and email addresses), and the firm’s data protection officer has flagged that the current manual process, which involves forwarding PDFs between personal email accounts and the operations lead’s inbox, does not meet the Article 5(1)(f) integrity and confidentiality requirement. The deadline is eight weeks: the founder wants a working pilot before the Q3 hiring decision, and the data protection officer wants a documented DPIA before any new system touches the invoice data.

    Approach: Two-Week Audit, Fixed-Scope Pilot, On-Premise Inference

    The engagement starts with a two-week AI automation audit. The team maps every document that enters the operations workflow, measures the current cycle time and error rate, and scores each workflow on volume, error cost, and automation feasibility. Invoice processing wins the composite score: 180 documents per month, a 18 percent error rate with a 20-to-35-minute correction cost per error, and a document format that maps cleanly to a structured extraction task. The pilot scope is fixed: extract vendor name, line items, tax codes, and total amount from PDF invoices, write the data to Xero via the API, and route flagged fields to the operations lead in Slack for approval. The architecture is model-agnostic: the orchestration service routes inference to an on-premise vLLM endpoint running a 7B-parameter open-weight model, because the GDPR review confirms that the invoice data cannot be sent to a cloud API. The Slack integration is built with the Slack Bolt framework, posting flagged items to a dedicated channel with approve and reject buttons. The human-in-the-loop gate is hard-coded: any field with a confidence score below 0.92 is flagged for human review.

    Outcome: 74 Percent Cycle-Time Reduction in Six Weeks

    The pilot runs for six weeks after the audit, with a two-week shadow period at the end where the AI drafts and the operations lead approves every output. The before/after baseline is measured over the final two weeks of the shadow run. The median cycle time drops from 4.2 days to 1.1 days, a 74 percent reduction. The manual correction rate falls from 18 percent to 4 percent. The operations lead reviews 22 flagged items per day in week one, dropping to 8 per day by week six as the model’s confidence improves on the firm’s specific vendor set. The senior operations lead, who had been spending roughly 14 hours per week on invoice processing, reports spending 3 hours per week on the approval queue and 2 hours per week on exception handling. The GDPR DPIA is completed in week three, documenting the data flows, the retention policy (invoices retained for seven years per UK tax law, extracted data retained for 12 months), and the human-in-the-loop approval gate. The on-premise hardware is a single workstation with an NVIDIA L40S 48 GB GPU, provisioned in week one and running the vLLM inference server for the duration of the pilot.

    Lessons for Similar Teams

    • The audit is the product, not the pilot. The two-week process audit produced a one-page baseline report that the client retained for internal reporting and the GDPR accountability record. The pilot was the validation, but the audit was the deliverable that justified the investment. Teams that skip the audit and jump straight to a pilot often discover mid-engagement that the workflow they chose is not the highest-impact one. – On-premise hardware is a procurement decision, not a technical one. The L40S workstation was ordered in week one, before the audit was complete. The lead time for GPU hardware in the UK is four to six weeks. Teams that order the hardware after the audit is done lose two to three weeks of the pilot timeline. – The Slack integration is the adoption lever. The operations lead approved 22 items per day in week one without any training, because the interface was the tool she already used. A separate dashboard would have added friction and likely reduced the approval rate below the threshold needed for the baseline comparison. – The confidence threshold is a tuning parameter, not a fixed constant. The 0.92 threshold for monetary fields was set in week one and adjusted to 0.95 in week four after the model’s performance on the firm’s specific vendor set improved. Teams that treat the threshold as a fixed constant either over-flag (wasting senior time) or under-flag (letting errors through). – The GDPR DPIA is a two-week task, not a one-day checkbox. The data protection officer spent three hours in week two reviewing the data flow diagram and two hours in week three reviewing the retention policy. The DPIA was completed in week three, not week one, because the model’s training data provenance had to be documented before the review could be signed off.
  • Deploying a GDPR-Compliant Voice Agent Over Confluence in 4 Weeks

    The Problem: Senior Staff Buried in Routine Knowledge Queries

    Your support team at a 201–500 person B2B SaaS company in the USA is drowning in repetitive internal knowledge queries. Senior engineers and support leads spend 30–40% of their week answering the same 20 questions about deployment procedures, API rate limits, and internal tooling, pulling them off the work that actually requires their judgment. You have already run isolated pilots on document extraction and invoice processing, but those pilots did not touch the voice channel or the internal knowledge base. The gap is specific: you need a voice agent that answers internal knowledge search queries from Confluence or Notion, built on LangChain and LangGraph, deployed in a 4-week integration sprint, and gated by GDPR compliance controls so that no personal data leaves the retrieval pipeline unreviewed. The goal is not to replace your support team; it is to free senior staff from routine work so they can focus on escalations, architecture decisions, and customer-facing strategy.

    Prerequisites Before the Sprint Starts

    Before the sprint starts, confirm the following are in place:

    • Confluence or Notion workspace access: a service account with read-only API tokens scoped to the specific spaces or databases the voice agent will index. For Confluence, this means a space-level API token; for Notion, an integration token with read permissions on the target databases.
    • Helpdesk staging environment: a sandbox instance of your ticketing system (Zendesk, Freshdesk, or Intercom) where the voice agent can be tested without affecting live customers.
    • 500+ historical tickets: exported as CSV with fields for query text, resolution, agent time, and category. This dataset builds the retrieval index and establishes the before/after baseline.
    • Compliance sign-off: a designated data protection officer or privacy counsel who has reviewed the Data Protection Impact Assessment (DPIA) and approved the lawful basis for processing under GDPR Article 6.
    • Voice infrastructure: API keys for a speech-to-text and text-to-speech provider (Twilio Voice, Amazon Polly, or Deepgram) and a webhook endpoint on your helpdesk to receive voice events.
    • LangGraph environment: a Python 3.11+ environment with langchain, langgraph, langchain-community, and your vector store driver (ChromaDB, Pinecone, or Weaviate) installed and tested locally.

    Step 1: Audit the Knowledge Base and Define the Query Taxonomy

    Spend the first five days mapping every internal knowledge query that reaches your support or engineering channels. Export 500 historical tickets from your helpdesk and tag each one with a category: deployment, API usage, internal tooling, billing, security, or other. Identify the top 15–20 categories that account for 70% of agent time. For each category, write a one-line description of the expected answer and note whether the answer contains personal data, contractual terms, or billing information. This last flag determines whether the query will route through the human-in-the-loop gate. Document the baseline: average cycle time per query (target: measure in minutes), error rate (percentage of answers that required correction), and the number of senior staff hours consumed per week. This baseline is the number you will compare against in week 4. Without it, you cannot prove the pilot delivered value.

    Step 2: Index Confluence Pages and Build the Retrieval Layer

    Build the retrieval pipeline in LangChain. Use the Confluence Cloud API (/wiki/rest/api/content) to pull page content as Markdown, strip HTML, and chunk the text into 512-token segments with 64-token overlap. Embed each chunk using text-embedding-3-small from OpenAI or a local nomic-embed-text model if data residency requires on-premises inference. Load the embeddings into a vector store (ChromaDB for a single-node pilot, Pinecone for multi-region). Write a Retriever class that accepts a query string, returns the top 5 chunks with similarity scores, and logs every retrieval hit. Before indexing, run a PII scanner over the corpus: flag any chunk containing email addresses, phone numbers, or names that match your customer database. If the PII hit rate exceeds 2%, pause indexing and add a redaction step that replaces flagged tokens with [REDACTED] before embedding. This step is non-negotiable under GDPR Article 5(1)(f), which requires integrity and confidentiality of personal data.

    Step 3: Build the LangGraph Voice-Agent Pipeline

    Define the LangGraph state machine with five nodes: intent_classification, retrieval, answer_synthesis, risk_gate, and voice_response. The intent_classification node uses a prompt that maps the user’s spoken query to one of your 15–20 categories and outputs a confidence score. If the score is below 0.7, the graph routes to a clarification node that asks the user to rephrase. The retrieval node calls the vector store and returns the top 5 chunks. The answer_synthesis node uses a system prompt that instructs the LLM to answer only from the retrieved context and to say “I don’t have that information” if the top similarity score is below 0.75. The risk_gate node checks whether the query category is flagged as high-risk (billing, security, personal data). If yes, the graph pauses and routes to a human approval queue via a Slack webhook or a simple web dashboard. The voice_response node sends the approved text to your TTS provider and streams the audio back to the caller. Each node’s state is serialized to a JSON file so the conversation can be resumed if the approval takes longer than 30 seconds.

    Step 4: Run the Pilot and Measure Before/After Baselines

    Run the pilot with a group of 10–15 internal users (support agents, junior engineers, and one senior lead) for five business days. Every interaction is logged: the raw audio, the transcribed query, the retrieved chunks, the similarity scores, the draft answer, the risk classification, the approval decision, and the final spoken response. At the end of the pilot, compute three metrics: cycle time (median seconds from query to spoken response, target: under 12 seconds for low-risk queries, under 45 seconds for high-risk queries with human approval), error rate (percentage of responses that the human reviewer edited or rejected, target: under 8%), and coverage (percentage of the 15–20 query categories that the agent answered without escalation, target: over 75%). Compare these numbers against the baseline from Step 1. If the error rate exceeds 15% or the cycle time for low-risk queries exceeds 20 seconds, do not proceed to rollout. Instead, tune the retrieval chunk size, adjust the similarity threshold, or add more few-shot examples to the answer_synthesis prompt. Document every tuning change in a changelog so the compliance team can audit the model’s behavior over time.

    Common Pitfalls and How to Detect Them

    Three failure modes will surface during the pilot, and each has a specific detection method. PII leakage in retrieval: the vector store returns a chunk containing a customer’s name or email, and the voice agent speaks it aloud. Detect this by running a PII scanner over every retrieval hit in the pilot logs and flagging any hit that returns a document with a flagged field. If the hit rate exceeds 2%, the indexing pipeline is leaking personal data. Hallucination on low-confidence retrieval: the agent generates an answer that is not supported by the retrieved context because the similarity score was just above the 0.75 threshold but the content was tangentially related. Detect this by logging the top-5 similarity scores for every query and flagging any response where the top score is between 0.75 and 0.85 for manual review. Approval queue bottleneck: the human-in-the-loop gate causes a 90-second delay because the reviewer is in a meeting. Detect this by measuring the median time from risk_gate entry to approval and alerting if it exceeds 30 seconds. If the bottleneck persists, add a second reviewer or a pre-approval rule for specific low-risk subcategories that do not require human sign-off.

  • AI Process Audit vs. Single-Process RAG Pilot: A Healthcare Company in Austria

    What Is Being Compared

    The two options are not alternatives in a vacuum; they are different scopes of the same engagement. Option A is a full AI process audit and roadmap: Forfis maps every back-office and customer-facing workflow, measures baseline cycle time and error rate on each, and produces a prioritised automation roadmap across the company. Option B is a single-process pilot: one workflow — here, an internal knowledge search assistant built on retrieval-augmented generation over the company’s Google Workspace documents — is scoped, built, and measured in a fixed three-month window. Both use the OpenAI API as the model layer, both integrate through existing APIs rather than replacing tools, and both ship with a human-in-the-loop approval gate. The difference is breadth: Option A covers the whole operation; Option B covers one process and proves the pattern before scaling.

    Criteria for the Comparison

    The judgment rests on seven criteria that matter to a 201-500 person healthcare company in Austria with no specific compliance mandate and a three-month timeline:

    • Time to first measurable value — how many weeks until a workflow runs with a before/after baseline.
    • Upfront cost — the fixed-scope fee for the audit or the pilot, before managed operation.
    • Breadth of coverage — how many workflows are mapped or automated by the end of the engagement.
    • Integration surface — which existing systems (Google Workspace, CRM, helpdesk) the AI layer touches.
    • Model-agnostic flexibility — whether the architecture can swap OpenAI for an open-weight model on client hardware if data-residency needs emerge.
    • Human-in-the-loop overhead — how many approval steps a support agent must complete per query.
    • Scalability path — how the engagement extends from one process to the next without re-scoping.

    Side-by-Side Comparison

    Criterion Option A: Full Audit + Roadmap Option B: Single-Process RAG Pilot
    Time to first measurable value 8-10 weeks (audit) + 4-6 weeks (first pilot) 3 weeks (audit slice) + 4-6 weeks (pilot)
    Upfront cost Higher: covers all workflows, multiple integrations Lower: one workflow, one integration (Google Workspace)
    Breadth of coverage All back-office and customer-facing workflows mapped One workflow: internal knowledge search
    Integration surface CRM, ERP, helpdesk, Google Workspace, messaging Google Workspace (Gmail, Drive, Calendar)
    Model-agnostic flexibility Full: per-workflow model selection Full: OpenAI API default, swappable
    Human-in-the-loop overhead Varies by workflow; set during audit Light: internal search, no money/health/contract decisions
    Scalability path Roadmap already built; next process is a scheduling decision Must re-scope for the second process

    When Each Option Wins

    Option B wins when the company’s immediate pain is concentrated in one workflow and the three-month timeline is a hard constraint. A 201-500 person healthcare company whose support team spends 25-40 minutes per ticket searching through Drive documents and Gmail threads will see a measurable cycle-time reduction within six weeks of the pilot starting. The RAG assistant indexes the existing Google Workspace content, retrieves the relevant SOP or device manual passage, and returns a grounded answer with a citation. The support agent approves the answer before sending it to the requester. No new hires are needed; the senior staff who previously handled routine knowledge lookups are freed to work on complex cases. The before/after baseline on time-to-answer and accuracy is captured in the first two weeks and compared at the end of the pilot.

    Option A wins when the company has multiple workflows with similar automation potential — invoice processing, document extraction, ticket triage, data entry — and the leadership team wants a single prioritised roadmap rather than a sequence of ad-hoc pilots. The audit maps all of them, measures baselines on each, and ranks them by expected cycle-time reduction and error-rate improvement. The cost is higher, but the company avoids the re-scoping overhead of going back to Forfis for every second process. For a company that has already automated one process and is now asking “what next?”, the audit is the natural next step.

    Recommendation for This Scenario

    For the scenario as specified — a 201-500 person healthcare and medtech company in Austria, no compliance mandate, three-month timeline, one process already automated, need to free senior staff from routine work, and a Google Workspace integration — Option B is the correct starting point. The company has already proven the pattern with one automated process; the next step is to apply the same pattern to internal knowledge search, not to commission a full audit that would extend the timeline beyond three months. The RAG pilot on Google Workspace is the highest-leverage single workflow for a support-heavy operation: it directly reduces the time senior staff spend on routine lookups, it integrates with the tools the team already uses, and it ships with a measured baseline that justifies the next investment. Once the pilot is live and the before/after numbers are in hand, the company can decide whether to commission the full audit (Option A) to map the remaining workflows, or to run a second pilot on a different process. The model-agnostic architecture means that if data-residency requirements emerge later, the OpenAI API layer can be swapped for an open-weight model on the company’s own hardware without re-architecting the integration.