Tag: Replace Manual Data Entry

  • AI Candidate Screening Pilot for a 2,000+ German Professional Services Firm

    The Problem: Manual Screening at Scale in a Regulated Environment

    You run a 2,000+ professional services firm in Germany. Your HR and recruiting team processes 15,000 to 40,000 applications per year across consulting, audit, and advisory practice areas. Each application requires manual data entry into your ATS, a screening pass against role-specific criteria, and a first-response email to the candidate. The cycle time from application receipt to first recruiter touch averages 3 to 5 business days. Your ISO 27001 certification requires documented controls over any system that processes candidate PII. You need to replace manual data entry, add round-the-clock candidate response, and scale the screening workflow across departments within 6 months. The constraint is fixed: a fixed-scope pilot on one workflow, measured against a before/after baseline, with human-in-the-loop approval for every classification that touches a candidate’s record.

    Prerequisites: What You Need Before Step 1

    Before you write a single line of integration code, confirm these items are in place:

    • Process audit completed. You have documented the 2 to 3 highest-volume screening workflows (e.g., junior analyst, associate, senior consultant) with their current cycle time, error rate, and volume. The audit identifies which fields are extracted manually and which classification rules recruiters apply.
    • Baseline measurement. You have measured cycle time and error rate on a sample of 200+ historical applications from the target workflow. This becomes your before/after benchmark.
    • ATS API access. Your ATS (Workday, SAP SuccessFactors, Taleo, or a German-specific system like Personio) exposes a REST API for candidate record updates and webhook endpoints for event notifications. You have API credentials and a sandbox environment.
    • ISO 27001 risk assessment. Your information security officer has documented a risk assessment for the AI component, covering data flow, PII handling, model output review, and rollback procedures.
    • Claude API access. You have an Anthropic API key with sufficient rate limits for the pilot volume. You have confirmed that candidate PII will be processed in EU data centers (Anthropic’s EU region) to satisfy GDPR and ISO 27001 data residency requirements.
    • Human-in-the-loop review dashboard. You have a simple interface where recruiters can approve, reject, or edit the AI’s classification before it writes to the ATS. This is non-negotiable under your ISO 27001 accountability controls.

    Step 1: Run the Process Audit and Measure the Baseline

    Run a structured process audit on the target workflow. Identify every manual step from application receipt to first recruiter touch. For each step, record: the input (PDF resume, email, form submission), the output (ATS record, classification tag, response email), the time spent, and the error rate. Use a sample of 200+ historical applications from the last 6 months. The audit output is a one-page workflow map with cycle time and error rate per step. This document becomes the baseline for your fixed-scope SOW. Without it, you cannot measure whether the pilot actually improved anything. The audit also identifies which fields are worth extracting: name, email, phone, location, years of experience, skill tags, education, and any role-specific criteria (e.g., ‘minimum 3 years in financial services’).

    Step 2: Define the Fixed-Scope Pilot SOW

    Define the fixed-scope SOW with your delivery partner. The SOW specifies: (1) which workflow is in scope (e.g., junior analyst screening), (2) which fields Claude extracts from the resume, (3) which classification rules apply (e.g., ‘meets minimum requirements: yes/no/partial’ based on years of experience and skill tags), (4) which human approval gates exist (every classification that writes to the ATS requires recruiter approval), (5) the integration points (ATS REST API, webhook endpoint, review dashboard), and (6) the success criteria (cycle time reduction target, error rate threshold, volume processed per week). The SOW is a fixed document. Any change after week 3 triggers a change request with a revised timeline and cost. This protects both parties from scope creep, which is the most common failure mode in AI pilots at 2,000+ firms.

    Step 3: Build the Claude API Extraction Pipeline

    Build the extraction pipeline. Your service receives the resume via a REST API endpoint (POST /api/v1/resumes) that accepts PDF or DOCX files. The service converts the document to text, then calls the Anthropic Claude API with a structured prompt that specifies the extraction schema. The prompt returns JSON with fields: name, email, phone, location, years_experience, skills (array), education, and a confidence score per field. The service validates the JSON schema, applies confidence thresholds (fields below 0.8 confidence are flagged for manual review), and stores the result in a temporary queue. The Claude API call uses the claude-sonnet-4-20250514 model for the balance of quality and cost. The prompt includes few-shot examples of correctly extracted resumes to reduce hallucination. The entire extraction takes 2 to 4 seconds per resume at the API level.

    Step 4: Implement Classification and Human-in-the-Loop Review

    After extraction, the service calls Claude a second time for classification. The prompt includes the extracted fields and the role-specific rubric (e.g., ‘Minimum 2 years experience in financial services, must hold a CFA charter or equivalent, fluent in German and English’). Claude returns a classification object: meets_requirements (boolean), confidence (float), summary (one-paragraph explanation), and flagged_fields (array of fields that triggered the classification). The service sends this classification to the human-in-the-loop review dashboard. The recruiter sees the extracted fields, the classification, and the summary. They can approve, reject, or edit before the classification writes to the ATS. The approval action triggers a webhook to your ATS endpoint (POST /api/v1/candidates/{id}/classification) with the final classification payload. The webhook uses HMAC-SHA256 signatures for authentication. This step ensures no AI classification touches a candidate’s record without human review, satisfying ISO 27001 accountability controls.

    Step 5: Integrate with Your ATS via REST API and Webhooks

    Integrate the screening service with your ATS via REST API and webhooks. The ATS sends new applications to your service via a webhook (POST /api/v1/webhooks/ats/application_received) with the candidate ID and document URL. Your service processes the resume, runs extraction and classification, and sends the result back to the ATS via a REST API call (PUT /api/v1/candidates/{id}). The ATS updates the candidate record with the extracted fields and classification. The webhook payload includes: candidate_id, extracted_fields (JSON), classification (JSON), metadata (model_version, processing_timestamp, source_document_hash). The webhook uses exponential backoff for retries (3 attempts, 1s/5s/30s delays). Log every webhook delivery with timestamp, payload hash, and response code. These logs become part of your ISO 27001 audit trail. The integration must handle edge cases: duplicate applications, malformed documents, and API rate limits from the ATS.

  • HIPAA-Compliant AI Invoice Processing for UK Healthcare: A Technical Deep Dive

    The Problem: Manual Invoice Processing in a Regulated Environment

    A 2,000+ employee healthcare organization in the UK processes 15,000 invoices monthly. Manual data entry takes 45 minutes per invoice, resulting in a 12-day average cycle time and a 3.2% error rate. The finance team spends 1,200 hours weekly on data entry, with 15% of time spent on error correction. The organization needs to reduce cycle time to under 48 hours and error rate to below 1% while maintaining HIPAA compliance. The challenge is not just automation but integration: the system must work with existing ERP (SAP S/4HANA), CRM (Salesforce), and helpdesk (Zendesk) without replacing them. The solution must handle complex invoice layouts, multi-currency transactions, and tax calculations while ensuring PHI never leaves the secure environment.

    The Mechanism: A Two-Stage Extraction Pipeline

    The pipeline uses a two-stage extraction. First, a vision-capable model (Claude 3.5 Sonnet) parses the PDF or image into structured JSON, identifying line items, totals, and vendor details. Second, a rule-based validation layer checks the JSON against the client’s chart of accounts and tax rules. If the confidence score drops below 0.85, the record is routed to a human reviewer. The system uses a hybrid approach: for high-volume, standardized invoices, a fine-tuned open-weight model (Llama 3 70B) runs on-premises. For complex, low-volume invoices, the system calls the Anthropic Claude API. The routing logic is based on invoice type, volume, and sensitivity. The pipeline exposes a /process-invoice endpoint that accepts PDFs and returns structured JSON. The ERP system calls this endpoint when a new invoice is uploaded. Conversely, the AI pipeline sends a webhook to the ERP when processing is complete, triggering automatic posting. For exceptions, the system sends a webhook to the client’s helpdesk, creating a ticket for human review.

    The Trade-offs: Accuracy, Cost, and Compliance

    The architect faces three key trade-offs. First, accuracy vs. cost: using the Claude API for all invoices costs $0.03 per invoice, while using an on-premises model costs $0.01 but requires $50,000 in hardware. The hybrid approach balances these costs. Second, compliance vs. flexibility: sending PHI to a third-party API violates HIPAA, but de-identifying data reduces accuracy. The solution is to send only financial metadata to the API, while patient identifiers remain in the client’s secure database. Third, speed vs. control: fully automated processing is faster but riskier. The human-in-the-loop approach adds 2-3 minutes per invoice but reduces error rates by 80%. The architect must also consider model drift: as invoice formats change, the model’s accuracy degrades. Retraining every 30 days mitigates this, but adds operational overhead. The managed operations model includes 24/7 monitoring, model retraining, and a dedicated support channel, covering these trade-offs.

    The Recommendation: A 3-Month Pilot with Managed Operations

    The pilot runs for 6-8 weeks. Week 1-2: process audit and data collection. Week 3-4: model fine-tuning and pipeline development. Week 5-6: parallel run (AI processes invoices alongside humans). Week 7-8: validation and go-live preparation. The 3-month timeline includes a 2-week buffer for stakeholder sign-off and integration testing with the ERP. The system tracks three key metrics: cycle time, error rate, and cost per invoice. Baselines are established during the process audit. During the pilot, the system compares AI performance against human performance. Post-implementation, the system monitors these metrics monthly and triggers retraining if error rates exceed 2% or cycle time increases by more than 10%. The managed operations model includes 24/7 monitoring, model retraining every 30 days, and a dedicated support channel. The client pays a monthly fee (typically 15-20% of the annual license cost) for ongoing optimization. This covers tracking model drift, updating validation rules, providing a monthly performance report, and handling API rate limits and cost optimization.

  • AI Invoice Processing for UK Professional Services: A 3-Month LangGraph Roadmap

    The Back-Office Bottleneck in UK Professional Services

    A 51-200 person professional services firm in the UK processes 800-1,500 invoices monthly. Each invoice requires manual data entry into the ERP, cross-referencing against purchase orders, and validation against vendor terms stored in Confluence or Notion. The baseline cycle time is 12-18 minutes per invoice, with a 3-5% error rate that triggers rework and payment delays. The operations team spends 40-60 hours weekly on this task, and the cost of errors (late payment penalties, vendor disputes) compounds over time.

    The problem is not a lack of tools. The firm already has an ERP, a helpdesk, and a knowledge base. The gap is in the workflow: data moves between systems through human hands, and each handoff introduces latency and error. AI workflow automation addresses this by replacing the manual extraction and validation steps with a model that reads the invoice, extracts fields, scores confidence, and routes exceptions to a human approver. The architecture plugs into existing systems via APIs rather than replacing them, preserving the firm’s current operational stack while automating the repetitive back-office work.

    LangGraph Stateful Workflow for Invoice Processing

    The system operates as a stateful graph defined in LangGraph. Each node represents a step: document ingestion, field extraction, validation, predictive scoring, and routing. The state object carries the invoice metadata, extracted fields, confidence scores, and approval status through the graph.

    [Ingest] → [Extract] → [Validate] → [Score] → [Route]
       ↑           ↑           ↑           ↑           ↓
       └───────────┴───────────┴───────────┴─────[Human Approve]
    

    The extraction node uses a vision-language model (GPT-4o or Claude 3.5 Sonnet) to parse the invoice PDF and output structured JSON. The validation node checks fields against the vendor master in the ERP and terms in Confluence/Notion via their APIs. The scoring node applies a predictive model that estimates the probability of payment delay or dispute based on historical data. If the confidence score falls below a threshold (typically 0.85), the graph routes to a human approval node where a person reviews the invoice and approves or rejects it. The approval action updates the state and triggers the next node, which posts the invoice to the ERP.

    The RAG layer indexes Confluence and Notion documents using semantic chunking (512-1024 tokens, 10-15% overlap) and stores embeddings in a vector store. At query time, the system retrieves relevant chunks on vendor terms, payment policies, and historical exceptions, augmenting the prompt to improve extraction accuracy.

    Trade-offs: Model Choice, Human-in-the-Loop, and Integration Depth

    The architect faces three key trade-offs. First, model choice: cloud APIs (OpenAI, Anthropic) offer higher quality but require data to leave the building, which conflicts with GDPR Article 22 if the data includes personal information. Open-weight models (Llama 3 70B, Mistral 7B) deployed on-premises via vLLM or TGI keep data local but require GPU infrastructure and yield slightly lower extraction accuracy. Forfis resolves this with a hybrid routing: invoices containing personal data go to the on-premises model; generic vendor data uses the cloud API.

    Second, human-in-the-loop granularity: a fully automated pipeline is faster but riskier. A fully manual approval is safe but defeats the purpose of automation. The compromise is confidence-based routing: only invoices below the threshold require human review. The threshold is tuned during the pilot to balance cycle time and error rate. A threshold of 0.85 typically routes 15-25% of invoices to humans, reducing manual work by 75-85% while keeping the error rate below 1%.

    Third, integration depth: shallow integration (API calls to ERP and helpdesk) is faster to deploy but misses opportunities for end-to-end automation. Deep integration (webhooks, event-driven updates) is more complex but enables real-time status tracking and audit trails. For a 3-month timeline, shallow integration is the pragmatic choice; deep integration can be added in a subsequent phase.

    3-Month Roadmap: Audit, Pilot, and Managed Operation

    For a 51-200 person UK professional services firm, the 3-month timeline breaks down as follows. Weeks 1-4: process audit and baseline measurement. The team maps the current invoice workflow, identifies the highest-volume and highest-error workflows, and measures cycle time and error rate. This baseline is critical for the before/after comparison that justifies the investment. Weeks 5-8: fixed-scope pilot on one workflow. The LangGraph workflow is deployed in a staging environment, and the team runs it on a sample of 100-200 invoices. The human-in-the-loop approval is tested, and the confidence threshold is tuned. Weeks 9-12: rollout and handover. The workflow is deployed to production, the dedicated AI team takes over managed operation, and the firm’s operations team is trained on the exception-handling dashboard.

    The dedicated AI team monitors key metrics: cycle time per invoice, error rate, human intervention rate, and model confidence distribution. If the error rate exceeds the baseline threshold, the team investigates whether the issue is in the extraction model, the validation rules, or the data quality. They also manage the RAG pipeline, re-indexing Confluence/Notion documents when content changes and monitoring retrieval accuracy. The service level agreement specifies 4-hour response times for production outages and weekly dashboards with monthly business reviews.

  • Fintech in the UAE: 8-Week Pilot to Automate Contract Review with On-Premise AI

    The 18-Minute Contract Review That Eats a Finance Team’s Week

    A 15-person fintech in the UAE processes 300 to 500 contracts per month. Each contract requires a finance analyst to open the document, locate the payment terms, extract the amounts, and enter them into SAP. The average cycle time is 18 minutes per contract, with a 7% error rate on data entry. The analyst spends 40% of their week on this task, which means they are not doing the reconciliation, forecasting, or vendor management that actually requires judgment. The pain is not that the work is hard; it is that it is repetitive, error-prone, and it consumes the time of the person who should be doing higher-value work. The metric that matters is not the cost of the analyst’s salary; it is the opportunity cost of the 40% of their week that is spent on data entry.

    Why Hiring More Analysts and Buying RPA Both Fail

    The first approach is to hire more analysts. This works until the volume grows, and then the problem scales with the headcount. The second approach is to use a commercial RPA tool to automate the data entry. RPA works for structured data in fixed formats, but contracts are semi-structured. The payment terms might be in a table, a paragraph, or a footnote. The RPA bot breaks when the format changes, and the maintenance cost of keeping the bot working across 500 different contract templates is higher than the cost of the analyst. The third approach is to use a commercial AI API to extract the data. This works, but the contract data leaves the building. For a fintech in the UAE, where the data includes payment terms, vendor names, and amounts, sending that data to a third-party API is a risk that the compliance team will flag. The problem is not that the technology is unavailable; it is that the available options do not fit the constraints of a small team with sensitive data and no dedicated compliance function.

    On-Premise RAG With a Human Approval Gate

    The approach that fits is a retrieval-augmented knowledge assistant built on open-weight models running on the company’s own hardware. The system ingests the contract, retrieves the relevant clauses, and extracts the payment terms, amounts, and dates. The output is a structured form that the finance analyst reviews and approves before it enters SAP. The model is model-agnostic: the pilot uses an open-weight model on-premise because the data cannot leave the building, but the architecture allows switching to a commercial API for workflows where the data is less sensitive. The integration is through the SAP API, not a replacement of SAP. The human-in-the-loop step is not a limitation; it is the design. The analyst sees the AI’s output, can edit it, and clicks approve. The system logs every approval and rejection, which creates an audit trail. The pilot is fixed-scope: one workflow, one integration, one measured baseline, 8 weeks.

    Eight Weeks From Audit to Measured Baseline

    Week 1: run the process audit. Map the contract review workflow step by step. Measure the current cycle time and error rate. Identify where the data enters and leaves the system. Check whether SAP has an API that can be used for integration. The output is a one-page recommendation with a projected ROI calculation. Week 2: select the model. For a fintech in the UAE where the data is sensitive, an open-weight model on the company’s own hardware is the right choice. The model should be capable of extracting structured data from semi-structured text. Week 3 to 4: build the RAG pipeline. Ingest the contract, retrieve the relevant clauses, extract the data, and populate the form. Week 5 to 6: build the approval interface. The analyst sees the AI’s output, can edit it, and clicks approve. The system logs every action. Week 7: integrate with SAP. The approved data enters the ERP through the API. Week 8: measure the baseline. Compare the cycle time and error rate against the pre-pilot numbers. The deliverable is a working system with documented metrics, not a proof of concept.

  • 4-Week AI Automation Pilot for Swiss Insurance Candidate Screening

    The Audit Phase: Mapping Manual Data Entry in Candidate Screening

    A 51-200 person insurance firm in Switzerland with no AI in production yet faces a specific problem: manual data entry in candidate screening, claims intake, and policy administration consumes 15-20 hours per week across three teams. The EU AI Act, which entered into force in August 2024, classifies candidate screening as a high-risk use case under Article 6(2), meaning you cannot simply deploy an AI model and walk away. You need a human-in-the-loop design, audit logs, and a measured baseline before you scale.

    The audit phase maps every step of the candidate screening workflow: resume ingestion, data extraction, classification against role requirements, drafting of initial assessments, and routing to a human reviewer. For a mid-size firm, this typically reveals that 60-80% of the time is spent on repetitive data entry and formatting, not on judgment. The audit output is a prioritized roadmap showing which workflow yields the highest ROI in the first 4-6 weeks.

    The key constraint is that the firm has no AI in production yet. This means the pilot must establish the baseline: cycle time per candidate, error rate on data entry, and time-to-first-response. Without this baseline, you cannot measure whether the automation actually works. The audit phase is not optional; it is the foundation for every subsequent decision.

    Building the Pilot: OpenAI API and Slack Integration

    The pilot uses the OpenAI API for drafting and classification tasks. For candidate screening, the model extracts structured data from resumes, classifies candidates against role requirements, and drafts an initial assessment. The orchestration layer plugs into the firm’s existing ATS via API, so the AI does not replace the system of record. Instead, it reduces manual data entry by 60-80% while keeping the human in the loop for final decisions.

    Integration with Slack or Microsoft Teams is critical for adoption. A recruiter receives a Slack message with the AI-drafted assessment and a one-click approve/reject button. This eliminates context switching and keeps the approval trail in a searchable channel. For a 51-200 person firm, this is the difference between a tool that gets used and one that sits in a dashboard nobody opens.

    The architecture is deliberately model-agnostic. If data residency rules change or the firm later needs to process health data, the orchestration layer stays the same while the model switches to an open-weight model on the client’s own hardware. This flexibility is not a nice-to-have; it is a requirement for a Swiss firm operating under the Federal Act on Data Protection (FADP) and the EU AI Act simultaneously.

    EU AI Act Compliance: Human Oversight and Audit Logs

    The EU AI Act requires you to document the AI system’s purpose, data sources, and human oversight mechanisms. For candidate screening, Article 14 mandates human oversight: the AI drafts, but a person approves. This is not a suggestion; it is a legal obligation. The firm must maintain a log of every AI-drafted assessment and the human’s decision, stored for at least 6 months and accessible to regulators on request.

    The pilot ships with a measured before/after baseline. Week 1 covers the process audit and baseline measurement. Weeks 2-3 build and test the automation with human-in-the-loop approval. Week 4 runs the pilot in production and measures cycle time and error rate against the baseline. Typical results show a 40-60% reduction in cycle time and a 30-50% drop in data entry errors for structured workflows.

    The compliance documentation is not a separate project; it is built into the pilot from day one. The audit trail, the human oversight log, and the baseline metrics are all part of the deliverable. This means the firm can demonstrate compliance to regulators without a separate documentation effort after the pilot ends.

    The 4-Week Timeline: Audit, Build, Measure

    The 4-week timeline is fixed-scope. Week 1: process audit and baseline measurement. The audit covers the candidate screening workflow end-to-end, identifying where manual data entry occurs and measuring cycle time and error rates. The output is a prioritized roadmap showing which steps to automate first.

    Weeks 2-3: build and test. The orchestration layer is configured to plug into the firm’s ATS via API. The OpenAI API is integrated for drafting and classification. The Slack or Microsoft Teams integration is tested with a small group of recruiters. The human-in-the-loop approval flow is validated: the AI drafts, the recruiter reviews, and the decision is logged.

    Week 4: production pilot and measurement. The workflow runs in production for one week. The firm measures cycle time per candidate, error rate on data entry, and time-to-first-response against the baseline. The deliverable is a before/after report with concrete numbers, not a qualitative summary. This report is the basis for the rollout decision and the managed operation pricing.

    Rollout and Managed Operation: What Comes After the Pilot

    The pilot is not the end; it is the proof point. After 4 weeks, the firm has a measured baseline, a working automation, and a compliance trail. The next step is rollout: extending the automation to other workflows, such as claims data entry or policy document extraction. The roadmap from the audit phase sequences these by ROI, starting with the workflow that has the clearest baseline and the least regulatory complexity.

    Managed operation is the ongoing service: monitoring the workflow, handling model updates, and maintaining the compliance documentation. For a 51-200 person firm, this is typically a monthly retainer of EUR 2,000-4,000, depending on the number of workflows and the volume of data processed. The retainer covers model monitoring, drift detection, and regulatory updates.

    The key lesson from the pilot is that the audit phase is not optional. Without a measured baseline, you cannot prove the automation works. Without a human-in-the-loop design, you cannot comply with the EU AI Act. Without a model-agnostic architecture, you cannot adapt to changing data residency rules. The 4-week pilot establishes all three, and the rollout builds on them.

  • 12-Point Checklist: Running a 4-Week AI Support Agent Pilot in Swiss Healthcare

    1. Run the process audit and lock the baseline

    Before writing a single line of prompt engineering, the audit must answer three questions: which workflow has the highest volume-to-complexity ratio, which data sources are API-accessible, and which compliance constraints are non-negotiable. For a Swiss healthcare company with no AI in production, the answer is usually ticket triage or first-response drafting on a customer support channel. The audit documents current cycle time (median minutes from ticket open to first human response) and error rate (misrouted or incomplete replies per 100 tickets). These two numbers become the baseline against which the pilot is measured. Without them, the pilot cannot prove ROI. The audit also maps every system the agent will touch—CRM, helpdesk, Notion or Confluence knowledge base—and confirms API credentials, rate limits, and data residency requirements. In Switzerland, FADP and the EU AI Act both apply; the audit flags which fields are personal data, which are health data, and which require human approval before any automated action. The output is a one-page roadmap: one workflow, one integration set, one success metric, four weeks. This document is the contract for the fixed-scope pilot and the reference for every subsequent decision.

    2. Define the fixed-scope pilot boundary

    The pilot scope must be narrow enough to finish in four weeks and broad enough to prove value. For a healthcare and medtech company, the typical scope is a conversational agent that triages incoming support tickets, drafts a first response using the company’s internal knowledge base, and routes the ticket to the right team. The agent does not close tickets, does not touch patient records, and does not send responses without human approval. The knowledge base lives in Notion or Confluence; the agent indexes those spaces via API and retrieves relevant passages to ground every draft. The CRM and helpdesk integrations are read-write for ticket metadata and read-only for customer history. The Anthropic Claude API handles classification and drafting; the model is selected for its instruction-following quality and context window, not for cost. The architecture is model-agnostic: if the client later moves to an open-weight model on local hardware for data residency reasons, the prompt layer and integration layer remain unchanged. The pilot ships with a dashboard showing cycle time, error rate, and human override rate, updated daily. At week four, the team compares the pilot numbers against the audit baseline and makes a go/no-go decision on rollout.

    3. Configure EU AI Act and Swiss FADP compliance gates

    The EU AI Act, effective in phases from 2025, requires transparency for AI systems that interact with humans. Article 50 mandates that users be informed they are interacting with an AI, unless it is obvious from context. For a healthcare support agent, this means the first message must state that the response is AI-drafted and subject to human review. The Act also classifies systems that make decisions affecting health as high-risk under Article 6, but a triage-and-draft agent that does not diagnose, prescribe, or alter treatment plans falls outside that category. Still, the agent must not process health data without a legal basis under GDPR and Swiss FADP. The pilot configuration includes a data classification layer: fields tagged as health data are routed to a human approver before any action. The agent’s system prompt explicitly forbids it from making medical claims, interpreting test results, or advising on treatment. Every response is logged with the model version, prompt hash, and retrieval context for auditability. The compliance checklist is signed off by the client’s data protection officer before the pilot goes live, and the log retention period matches the client’s regulatory requirement, typically 12 months for healthcare records in Switzerland.

    4. Build the retrieval layer over Notion or Confluence

    The agent’s value depends on retrieval quality. The knowledge base in Notion or Confluence must be structured so the agent can find the right passage in under 200 ms. Before the pilot, the team runs a retrieval audit: take 50 real support tickets from the past quarter, identify the correct knowledge base article for each, and measure how often a vector search over the raw document text returns that article in the top three results. If the hit rate is below 80%, the knowledge base needs restructuring before the agent is built. Concretely, this means splitting long pages into discrete, self-contained sections, adding metadata tags (product, issue type, severity), and removing deprecated content. The retrieval pipeline uses a hybrid approach: dense vector embeddings for semantic matching and BM25 for exact keyword hits, with a reranking step using the Claude API to score the top ten candidates. The agent’s system prompt instructs it to cite the specific knowledge base section in every draft, so the human approver can verify the source. If the retrieval confidence score falls below a threshold the team sets during the audit, the agent flags the ticket for manual handling rather than drafting a potentially wrong response. This guardrail is non-negotiable in a healthcare context.

    5. Measure cycle time, error rate, and override rate daily

    The pilot runs for four weeks with a daily standup and a weekly metrics review. The team tracks three numbers every day: median cycle time from ticket open to first human-approved response, error rate (tickets requiring rework after approval), and human override rate (percentage of drafts the approver rejects or significantly edits). The audit baseline from step one is the reference. A successful pilot shows at least a 30% reduction in cycle time and a 20% reduction in error rate, with an override rate below 15% by week three. If the override rate stays above 25%, the team investigates: is the retrieval missing the right article, is the prompt too vague, or is the knowledge base outdated? The fix is applied within 48 hours and the metrics are re-measured. The pilot also includes a shadow mode for the first three days: the agent drafts responses but does not send them; the human approver compares the draft against what they would have written. This calibrates the prompt and the retrieval thresholds before the agent goes live. At the end of week four, the team produces a one-page report: baseline vs. pilot numbers, override rate trend, top five failure modes, and a recommendation on rollout scope. The report is the input to the next engagement, not a marketing document.

    6. Maintain the checklist and the agent after go-live

    The pilot is not a one-and-done deliverable. The knowledge base in Notion or Confluence changes weekly; new product releases, policy updates, and support macros all alter the retrieval landscape. The team schedules a monthly retrieval audit: take 20 new tickets, measure the hit rate, and restructure sections if the rate drops below 80%. The prompt layer is versioned in a repository with a changelog; every change is tested against a fixed set of 30 evaluation tickets before deployment. The compliance log is reviewed quarterly by the data protection officer to confirm that no health data was processed without approval and that the AI transparency notice is still present in every first response. The model provider’s terms of service and the EU AI Act’s obligations are re-checked at each quarterly review, because both evolve. The team also maintains a runbook for model degradation: if the Claude API’s response quality drops due to a provider-side change, the runbook specifies the fallback—switch to the open-weight model on local hardware, re-run the evaluation set, and deploy within 24 hours. The checklist itself is stored in the same Notion or Confluence space the agent indexes, so the team can search for it the same way the agent searches for support articles. This keeps the maintenance process visible and auditable.

  • LangGraph AI Agent for HR Workflow Orchestration in an Austrian Fintech

    The Problem: Fragmented HR Data Entry in a 30-Person Austrian Fintech

    A 30-person fintech in Vienna processes 40-60 onboarding documents per month: contracts, bank details, compliance attestations, and internal policy acknowledgments. Each document requires a human to extract fields, cross-reference against the HR system, and log the data into three separate tools. The median cycle time is 72 minutes per document, and the error rate on manual data entry sits at 4-6%, triggering rework and compliance risk under GDPR Article 5(1)(d) (accuracy of personal data). The problem is not volume but fragmentation: the data lives in PDFs, email threads, and a legacy HR system, and no single tool connects them. The automation target is not to replace the HR team but to eliminate the 12-15 hours per week of manual data entry and document routing that currently consume senior staff time. The constraint is strict: personal data cannot leave Austrian or EU jurisdiction, and any automated action affecting a candidate or employee requires human approval under GDPR Article 22.

    Mechanism: LangGraph State Machine and RAG Pipeline

    The architecture uses LangGraph as the orchestration layer and LangChain for LLM and vector store abstractions. LangGraph models the workflow as a stateful directed graph with nodes for intake, classification, RAG retrieval, draft generation, human approval, and dispatch. Each node is a Python function that receives and returns a state object. The graph supports conditional edges: if the classifier flags a document as high-risk (e.g., a contract amendment), the path routes to a senior reviewer; if it is a routine bank-detail update, it routes to a junior approver. The state persists in PostgreSQL via LangGraph’s checkpoint store, so the workflow survives process restarts. The RAG pipeline ingests internal policy docs, onboarding checklists, and HR system exports. Documents are chunked at 512 tokens with 64-token overlap, embedded using BGE-M3 (multilingual, supports German and English), and stored in pgvector. At query time, the agent retrieves the top-5 chunks, constructs a context-augmented prompt, and generates a structured JSON response with extracted fields and a confidence score. The LLM layer is model-agnostic: OpenAI GPT-4o handles general knowledge queries where no personal data is in the prompt, while Llama 3 70B running on the client’s own GPU server handles any task involving personal data, ensuring GDPR data residency.

    Trade-offs: Model Choice, Approval Granularity, and Integration Depth

    Three architectural choices dominate the trade-off space. First, model selection: using OpenAI or Anthropic APIs reduces infrastructure cost and improves quality on complex reasoning, but personal data in the prompt violates GDPR data residency for an Austrian company. The cost of using open-weight models on client hardware is a 15-20% drop in classification accuracy on edge cases and a one-time GPU server cost of EUR 8,000-12,000. Second, human-in-the-loop granularity: inserting an approval node after every agent action maximizes compliance but adds 5-10 minutes of latency per document. A tiered approach, where routine documents auto-approve after a 24-hour window and high-risk documents require immediate human review, reduces latency by 40% but requires a well-defined risk taxonomy. Third, integration depth: building a custom UI for HR staff gives full control but adds 2-3 weeks of development. Integrating with Slack or Microsoft Teams via their existing APIs (Slack Block Kit, Teams Adaptive Cards) reuses the tools the team already uses, cuts development time by 60%, and keeps the approval workflow in the channel where the document was originally shared. The Teams integration uses the Bot Framework with a webhook endpoint; the Slack integration uses a slash command that triggers the LangGraph agent via a REST API.

    Recommendation: 8-Week Integration Sprint for One Process

    For a 30-person Austrian fintech, the 8-week sprint follows a fixed sequence. Weeks 1-2: process audit. Map every HR document type, identify the three highest-volume workflows (typically onboarding data entry, policy acknowledgment tracking, and candidate status updates), and measure baseline cycle time and error rate. Weeks 3-4: build the LangGraph agent. Scaffold the state machine, implement the RAG pipeline, and connect to the HR system API. Deploy the open-weight model on the client’s hardware. Weeks 5-6: integrate with Slack or Teams. Build the interactive approval cards, test the webhook flow, and configure the checkpoint store. Weeks 7-8: pilot and measure. Run the agent on one workflow (e.g., onboarding document processing) for two weeks, with a human approving every action. Measure cycle time, error rate, and manual hours saved against the baseline. The pilot ships with a before/after report. The recommendation is to start with the workflow that has the highest volume and the lowest compliance risk, not the most complex one. For a fintech, that is usually routine onboarding data entry, not contract amendment review. The agent should be scoped to extract and classify, not to make decisions. Every output that touches a candidate’s or employee’s data must pass through a human approval node before it is written to the HR system or sent to the individual.

  • UK Fintech AI Data Enrichment Pilot: 4-Week ISO 27001-Compliant Automation

    The Problem: Manual Data Entry in a Regulated Fintech

    A 51-200 person UK fintech running ISO 27001 faces a specific constraint: compliance data cannot leave the building, yet the team is drowning in manual data entry for client onboarding, transaction enrichment, and regulatory reporting. The process audit identifies one workflow—say, enriching client records from source documents into the CRM—where cycle time is 14 minutes per record and error rate sits at 3.2%. The fixed-scope pilot targets that single process, with a four-week timeline and a measured before/after baseline on both metrics.

    The architecture is deliberately model-agnostic. Open-weight models run on the client’s own hardware, satisfying ISO 27001 Annex A.13 and A.14 requirements without relying on third-party API providers. The AI layer drafts the enriched data, a human reviewer approves anything touching compliance records, and the final output is written back to the existing CRM via its API. No new software is installed; the integration plugs into the system the team already runs.

    The Four-Week Pilot: Audit, Build, Measure

    Week 1 covers the process audit and baseline measurement. The team documents the current workflow: where records originate, which fields are manually entered, where errors occur, and what the cycle time is per record. A sample of 50 records is processed manually to establish the baseline: 14 minutes average cycle time, 3.2% error rate.

    Weeks 2 and 3 cover model configuration and integration. The open-weight model is fine-tuned or prompted to extract and enrich the specific fields in the target workflow. The integration is built through the CRM’s API, so the enriched data lands in the same system the team already uses. Slack or Microsoft Teams is connected via its API, so the human reviewer receives AI-drafted enrichments in the channel they already use, approves or edits them, and the final record is written back.

    Week 4 covers human-in-the-loop testing and final metrics. The same 50-record sample is processed through the automated workflow. The before/after report documents cycle time, error rate, and the number of records requiring human intervention. The pilot ends with a documented deliverable, not an open-ended deployment.

    Compliance: ISO 27001 and On-Premise Models

    ISO 27001 requires documented risk assessment, access control, and audit trails for all systems handling sensitive data. An on-premise open-weight model satisfies the data residency and access control requirements because regulated data never leaves the client’s hardware. The human-in-the-loop approval step provides the audit trail that ISO 27001 Annex A.12.4 (logging and monitoring) expects for automated decisions affecting compliance records.

    The model-agnostic architecture means the company is not locked into a single vendor. If the open-weight model’s quality is insufficient for a specific task, the architecture can route that task to a hosted API where data can leave the building. For a UK fintech with ISO 27001 obligations, the on-premise option is the default for compliance-sensitive workflows, but the architecture allows flexibility where the risk profile permits.

    The integration with Slack or Microsoft Teams keeps the workflow within the team’s existing communication pattern. No new software is installed, no new training is required beyond the approval step, and the audit trail is logged in the same channel the team already uses.

    Scaling Without New Hires: The Operational Payoff

    The pilot replaces manual data entry by extracting, validating, and enriching records from source documents or systems. The AI layer drafts the enriched data, a human reviewer approves anything touching compliance or financial records, and the final output is written back to the existing CRM or ERP via its API. The before/after baseline measures cycle time and error rate on the same sample of records, so the improvement is quantified, not assumed.

    For a 51-200 person company, the goal is to scale operations without new hires. The AI handles the repetitive extraction and enrichment, freeing the team to focus on judgment calls and exceptions. The fixed-scope structure means the pilot ends with measured metrics, not an open-ended deployment. The company then decides whether to scale to additional workflows based on the documented before/after report.

    The internal knowledge search assistant is a natural extension of the same architecture. It uses retrieval-augmented generation over the company’s own documentation, CRM records, and compliance policies. The AI retrieves relevant passages and drafts a response, which a human reviewer can approve or edit before it is shared. This replaces the manual process of searching through PDFs, shared drives, and CRM notes to answer internal queries.

  • Austrian Insurtech Cuts Support Cycle Time 50% with Voice Agent and RAG Pilot

    Background: A 300-Person Austrian Insurtech

    This case study is a composite based on patterns observed across multiple engagements. We do not name real customers. The company described here matches the profile of a mid-sized Austrian insurtech: 300 employees, 12 years in operation, serving private and small-business customers across Austria and Germany. The stack includes a legacy CRM (Salesforce), an ERP (SAP), and a helpdesk (Zendesk). Internal documentation lives in Confluence, with some policy procedures in Notion. The company had been using basic rule-based chatbots for two years but had not moved to generative AI. The operations team was under pressure to scale support without adding headcount, as the Austrian labor market for customer support specialists was tight and salaries had risen 12% year-over-year.

    Challenge: Scaling Support Without New Hires

    The operations director identified three specific pain points. First, 45% of inbound support tickets involved repetitive data entry: policy number lookups, claim status updates, and address changes. Second, agents spent an average of 14 minutes per ticket searching internal documentation for policy details and claim procedures. Third, the company faced a compliance deadline under the EU AI Act, which required transparency and human oversight for customer-facing AI systems. The deadline was 18 months out, but the company wanted to be ahead of the curve. The operations team had 12 full-time support agents, and the director was told by HR that hiring two more would cost EUR 120,000 annually. The goal was to replace manual data entry and reduce documentation search time without adding headcount.

    Approach: Process Audit and Fixed-Scope Pilot

    The engagement began with a two-week process audit. We mapped every step of the top 20 support workflows, measured cycle time and error rate for each, and identified where manual data entry occurred. The audit revealed that 60% of the top 20 workflows involved repetitive data entry that could be automated. We then built a fixed-scope pilot targeting one workflow: first-response triage for policy status inquiries. The pilot used Anthropic Claude API for the voice agent, with a RAG assistant indexing Confluence and Notion documentation. The architecture was model-agnostic, so we could switch to an open-weight model on the client’s hardware if data residency became an issue. The pilot integrated with Salesforce and Zendesk through their APIs, not by replacing them. Human-in-the-loop approval was built in: the voice agent drafted responses and extracted data fields, but a human approved anything that touched money, health data, or a contract.

    Outcome: Measured Baseline and Rollout Decision

    The 8-week pilot delivered measurable results. Average ticket resolution time for policy status inquiries dropped from 14 minutes to 7 minutes, a 50% reduction. Manual data entry errors fell from 8% to 2%, a 75% reduction. The voice agent handled first-response triage for 70% of policy status inquiries, reducing the need for human escalation. The RAG assistant cut documentation search time from 14 minutes to 3 minutes per ticket. The human-in-the-loop approval process added 2 minutes to each ticket, but the net effect was a 5-minute reduction in cycle time. The pilot met the EU AI Act transparency requirements: all interactions were logged, and the voice agent disclosed its AI nature to customers. The operations director approved a rollout to the remaining 19 workflows, with a target of 12 months for full deployment.

    Lessons for Similar Teams

    • Start with the process audit, not the model. The audit revealed that 60% of the top 20 workflows were automatable, but the model choice was secondary. Teams that skip the audit and jump to model selection often automate the wrong workflows.
    • Fixed-scope pilots reduce risk. The 8-week timeline and defined success metrics gave the operations director confidence to approve the rollout. Without the pilot, the rollout would have been a 6-month project with no baseline to measure against.
    • Human-in-the-loop is not optional. The EU AI Act requires human oversight for customer-facing AI systems. Building it in from the start avoids rework and reduces liability risk.
    • Model-agnostic architecture future-proofs the investment. The ability to switch between Anthropic Claude and open-weight models on the client’s hardware means the company can adapt to changes in cost, latency, and compliance requirements without rebuilding the system.
    • Integrate with existing systems, not replace them. The pilot plugged into Salesforce, Zendesk, and Confluence through their APIs. This reduced integration risk and allowed the operations team to continue using the tools they already knew.
  • LLM Document Extraction with n8n: EU AI Act Compliance for B2B SaaS in Austria

    EU AI Act

    The EU AI Act (Regulation (EU) 2024/1689) is the first comprehensive AI regulation in the world, entering into force on 1 August 2024. It classifies AI systems by risk level and imposes obligations on providers and deployers. For a document extraction pipeline that processes order and shipment data, the system is generally not high-risk, but if it touches personal data or feeds automated decisions, it may trigger transparency and logging obligations under Articles 13 and 14. The Act’s Article 4 requires AI literacy for staff operating the system, which Forfis addresses through the pilot’s training module. In this scenario, the compliance checklist maps each pipeline step to the relevant Act articles, ensuring the client can demonstrate conformity during audits.

    Document Extraction

    Document extraction is the process of converting unstructured or semi-structured documents (PDFs, emails, scanned images) into structured data (JSON, CSV, database records). In this scenario, the LLM reads order confirmations and shipment notifications from Gmail, extracts fields like order ID, shipment ID, carrier, and tracking number, and outputs them as JSON. The extraction accuracy depends on the document format and the LLM’s training data; Forfis measures accuracy per field during the pilot and reports it in the baseline. The human-in-the-loop review step catches extraction errors before the data is written to the SaaS platform, reducing the error rate to below 0.5% in Forfis’s measured baselines.

    Human-in-the-loop (HITL)

    Human-in-the-loop (HITL) means a person reviews and approves the AI’s output before it affects downstream systems. In this pipeline, the LLM extracts order and shipment data, but a human operator confirms the extracted fields before the data is written to the B2B SaaS platform. This is mandatory under Forfis’s default delivery model for anything touching financial records or customer commitments. The HITL step adds roughly 30–60 seconds per document but reduces error rates to below 0.5% in Forfis’s measured baselines. The EU AI Act’s Article 14 requires human oversight for high-risk systems, and the HITL review step satisfies this requirement by allowing the operator to reject, correct, or escalate the extracted data.

    n8n Orchestration

    n8n is an open-source workflow automation platform that uses a visual node-based editor to connect APIs, databases, and services. In this scenario, n8n acts as the orchestration layer: it receives a new email from Google Workspace, triggers the LLM extraction node, validates the output against a schema, and pushes the structured data into the B2B SaaS platform’s order management API. n8n’s self-hosted deployment option keeps data within the client’s Austrian infrastructure, satisfying data residency requirements. The platform’s node-based architecture means the pipeline can be modified without code changes, and the model-agnostic design allows swapping between OpenAI, Anthropic, or open-weight models by changing a single configuration parameter.

    LLM Integration

    LLM integration refers to embedding a large language model into an existing system to perform a specific task, such as document extraction or text classification. In this scenario, the LLM is integrated into the n8n pipeline to read order and shipment emails and extract structured data. The integration is model-agnostic: Forfis uses OpenAI’s GPT-4o or Anthropic’s Claude 3.5 Sonnet for cloud-based processing, or an open-weight model like Llama 3 70B on the client’s own GPU server for regulated data. The n8n orchestration layer abstracts the model choice, so switching providers requires only a configuration change, not a code rewrite. The LLM’s output is validated against a JSON schema before being pushed to the SaaS platform.

    Fixed-Scope Pilot

    Fixed-scope pilot is a bounded engagement with a defined deliverable, timeline, and success metric. Here, the pilot runs for two weeks, targets one specific workflow (order and shipment status updates), and ships with a measured before/after baseline on cycle time and error rate. The scope excludes multi-language support, voice interfaces, or integration with systems outside the agreed API list. This structure limits risk for the client and gives Forfis a clear acceptance criterion. The pilot report compares the baseline metrics from the first three days (manual process) with the metrics from the remaining nine days (automated pipeline), quantifying the reduction in cycle time and error rate as the business case for full rollout.

    Process Audit

    Process audit is the first phase of Forfis’s delivery model, typically taking two to three days. Forfis interviews the operations team, observes the current manual workflow, and maps every step from email receipt to data entry completion. The audit identifies which fields are extracted, which systems are involved, where errors occur, and how long each step takes. The output is a process map and a recommendation on which workflow to automate first. In this scenario, the audit confirmed that order and shipment status updates were the highest-volume, most error-prone workflow, making it the ideal pilot candidate. The audit also identifies compliance requirements under the EU AI Act and data residency constraints that shape the architecture.