Blog

  • 12-Point Checklist: AI Order Status Automation for Swiss Professional Services

    1. Verify the workflow scope and baseline metrics

    Before writing a single line of code, confirm the workflow you are automating is the right one. For a 51-200 person professional services firm in Switzerland, order and shipment status updates in customer support typically consume 15-25% of agent time. Verify that the ticket volume justifies automation: if fewer than 200 tickets per month require status lookups, the ROI may not support the integration cost. Document the current process: how an agent receives a status inquiry, which system they check (ERP, logistics portal, email chain), how long the lookup takes, and what format the response takes. This baseline becomes the denominator for your before/after measurement. Without it, you cannot prove the pilot delivered value. The audit should also flag any tickets that involve personal data under GDPR, because those will need a different handling path than purely transactional status queries.

    2. Document the GDPR and Swiss FADP compliance path

    GDPR and the revised Swiss FADP (effective 1 September 2023) require a documented legal basis for processing personal data. For order status updates, the data typically includes customer name, email, order ID, and shipment tracking number. Confirm that your privacy notice covers automated processing of this data. If the predictive scoring model uses customer history to estimate resolution time, you need a legitimate interest assessment under GDPR Article 6(1)(f) or explicit consent under Article 6(1)(a). Log every model inference: timestamp, input data, model version, prompt, and output. Store these logs for at least 6 months to support data subject access requests under GDPR Article 15. Assign a data protection officer or responsible person to review the processing record. If any data leaves Switzerland, ensure a standard contractual clause or adequacy decision covers the transfer, even if the data is pseudonymized.

    3. Configure the OpenAI API endpoint and prompt constraints

    Provision the OpenAI API key in a secrets manager, not in code. Use GPT-4o-mini for cost efficiency on high-volume status lookups; reserve GPT-4o for complex edge cases where the model must interpret ambiguous shipment data. Set the temperature parameter to 0.1 for deterministic output. Write a system prompt that constrains the model to factual status language: “You are a customer support assistant. Respond only with the order status, expected delivery date, and any delay reason. Do not speculate. If the data is missing, state that clearly.” Test the prompt with 20 real ticket samples from the past month. Measure accuracy: the model should correctly state the status in at least 90% of cases before you move to integration. Log token usage per request to forecast monthly API costs. For a firm processing 5,000 tickets per month, expect roughly CHF 50-150 in API costs at GPT-4o-mini rates.

    4. Integrate with Zendesk or Intercom via webhooks and REST APIs

    Subscribe to the ticket.created and ticket.updated webhooks in Zendesk or Intercom. In Zendesk, create a trigger that fires when a ticket is tagged “status-inquiry” and routes it to your automation endpoint. In Intercom, use the webhook for new conversations and filter by custom attributes. The automation layer receives the ticket ID, customer email, and message body. It queries the order management system via API for the current status, passes the result to the LLM, and posts the response back through the helpdesk API. Handle rate limits explicitly: Zendesk allows 200 requests per minute per user, Intercom allows 100. Implement exponential backoff for 429 responses. Test the full loop with 10 real tickets in a staging environment before touching production. Verify that the response appears in the correct ticket thread and that the agent can see the AI-generated draft before it is sent.

    5. Implement the human-in-the-loop approval gate

    The model drafts the status response; a human approves it before it reaches the client. This is non-negotiable for GDPR compliance and for maintaining trust in a professional services context. Configure the helpdesk to flag AI-generated responses with a visible indicator. The agent reviews the draft, checks it against the order data, and either sends it as-is or edits it. Log every approval, edit, and rejection. This log serves two purposes: it provides an audit trail for GDPR Article 30 records of processing, and it gives you training data to improve the prompt over time. If the agent rejects the AI response more than 10% of the time in the first two weeks, pause the automation and revisit the prompt or the data source. The human-in-the-loop step should add no more than 30 seconds to the agent’s workflow; if it takes longer, the integration is not working correctly.

    6. Automate the monthly reporting pipeline

    Automate the data collection for monthly reporting, but keep the narrative summary human-written for the first three months. The report should include: total tickets processed, percentage handled by AI vs. human, average cycle time before and after automation, error rate (incorrect or incomplete status updates), escalation rate, and customer satisfaction scores from post-interaction surveys. Store the raw data in a simple database or a structured spreadsheet. Generate the report on the 1st of each month and send it to stakeholders as a one-page PDF with two charts: cycle time trend and error rate trend. The before/after baseline must use the same ticket categories and the same measurement method. If the AI reduces cycle time from 4.2 minutes to 1.1 minutes and cuts error rate from 8% to 2%, that is your ROI story. Automate the data pull; do not automate the interpretation until the data is stable for at least three months.

    7. Maintain the checklist as a living document

    The checklist is a living document, not a one-time artifact. Review it after each sprint and after any significant change: a new model version, a change in ticket volume, a regulatory update, or a shift in the order management system. Assign a single owner for the checklist, typically the technical lead on the engagement. Update it within 48 hours of any change that affects the automation. Archive old versions with a date stamp so you can trace what was in place when a specific incident occurred. If the firm adds a new use case, such as invoice processing or document extraction, create a separate checklist for that workflow rather than bloating this one. The checklist should remain under 20 items; if it grows beyond that, split it into sub-checklists by function. Re-validate the GDPR compliance section quarterly, because data protection regulations in Switzerland and the EU are actively evolving, and the FADP enforcement guidance from the FDPIC is updated regularly.

  • Ticket Triage and Routing for a 51-200 Person B2B SaaS Company: A Two-Week Pilot

    The problem: manual triage across three languages

    Your support team handles 150 to 400 tickets per day across English, Spanish, and German. Each ticket is read, categorized, and routed by a human agent before any response is drafted. The median cycle time from ticket creation to first human action is 42 minutes. You want to cut that number without adding headcount, and you want the routing to work across all three languages without a separate team per locale. The constraint is that you cannot replace your helpdesk or CRM. The model must plug into the REST API and webhook endpoints you already expose, and the pilot must be scoped so that you know the total cost and the success criteria before the first sprint starts.

    Prerequisites before step 1

    Before the first sprint, confirm the following are in place:

    • Helpdesk API access. A service account with read and write permissions on ticket objects. The account must be able to create, update, and query tickets via REST. Verify that the API rate limit is at least 100 requests per minute.
    • Webhook endpoint. A publicly reachable HTTPS URL that accepts POST requests with a JSON body. The endpoint must return a 200 status within 5 seconds. If your helpdesk does not natively support webhooks, you will need a lightweight relay service.
    • Historical ticket data. At least 500 labeled tickets per language, exported as CSV or JSON. Each record must include the ticket body, the final category, the assigned team, and the language tag. This is the training set for the scoring model.
    • OpenAI API key. A key with access to the GPT-4o or GPT-4o-mini model. The key must have sufficient credits for the pilot volume. For 300 tickets per day over 14 days, budget for roughly 4,200 API calls.
    • A named owner. One person on your side who can approve scope changes, answer integration questions, and sign off on the pilot results. This person should have authority over the helpdesk configuration.

    Step 1: Export and label your historical tickets

    Export 500 to 1,000 tickets per language from your helpdesk. Each record must contain the ticket body, the final category assigned by a human, the team that handled it, and the language tag. If your helpdesk does not store a language tag, infer it from the ticket body using a language-detection library such as langdetect or fasttext. Save the export as tickets_train.csv with columns: ticket_id, body, category, team, language. Split the file into a 70% training set and a 30% validation set. The validation set is used to measure routing accuracy before the model goes live. If any category has fewer than 50 examples, merge it with a related category or flag it for manual review in the pilot.

    Step 2: Configure the OpenAI scoring model

    Build a scoring function that takes a ticket body and returns a category label, a confidence score from 0 to 100, and a language tag. Use the OpenAI API with the GPT-4o model. The prompt should include the list of valid categories, the language of the ticket, and the instruction to return JSON with fields category, confidence, and language. Set the temperature parameter to 0.1 to reduce variance. Set max_tokens to 200. The function should handle API errors by retrying up to three times with exponential backoff (1 second, 5 seconds, 25 seconds). If all three retries fail, return a default category of unclassified with a confidence score of 0. Log every API call with the ticket ID, the model version, and the latency in milliseconds. Store the logs in a file or a lightweight database for the pilot review.

    Step 3: Build the webhook-to-helpdesk router

    Write a webhook handler that receives the scoring result and calls your helpdesk REST API to update the ticket’s routing field. The handler should accept a POST request with a JSON body containing ticket_id, category, confidence, and language. It should call the helpdesk API endpoint PATCH /tickets/{ticket_id} with a JSON body that sets the routing field to the predicted category and the priority field based on the confidence score. If the confidence score is 85 or above, set the priority to auto. If the score is between 60 and 84, set the priority to review. If the score is below 60, set the priority to manual. The handler must return a 200 status to the caller within 5 seconds. If the helpdesk API returns an error, log the error and retry up to three times. After the third failure, write the ticket ID to a dead-letter queue file.

    Step 4: Validate routing accuracy on the holdout set

    Run the scoring model on the 30% validation set from step 1. For each ticket, compare the predicted category to the human-assigned category. Calculate the routing accuracy as the percentage of tickets where the predicted category matches the human category. Calculate the median confidence score for correctly routed tickets and for incorrectly routed tickets. If the routing accuracy is below 80%, review the misclassified tickets and adjust the prompt or the category definitions. If the median confidence for correct tickets is below 70, lower the confidence threshold for auto-routing. Document the final thresholds in a configuration file named triage_config.json with fields auto_threshold, review_threshold, and manual_threshold. This file is read by the webhook handler at startup.

    Step 5: Run the two-week pilot

    Deploy the webhook handler to a staging environment that mirrors your production helpdesk configuration. Send 50 test tickets through the full pipeline: ticket creation in the helpdesk, webhook trigger, scoring model call, routing update. Verify that each ticket is routed to the correct team and that the priority field is set according to the confidence thresholds. Check the dead-letter queue file for any failed deliveries. Monitor the API latency for each scoring call. The median latency should be under 800 milliseconds. If the median latency exceeds 1,200 milliseconds, reduce the max_tokens parameter or switch to the GPT-4o-mini model. Once all 50 test tickets pass, promote the handler to production and enable the webhook on your live helpdesk instance.

  • LangGraph Agent vs. Managed Pilot: HR Back-Office Automation in Swiss Healthcare

    What Is Being Compared: In-House LangGraph Agent vs. Managed Fixed-Scope Pilot

    The two options under evaluation are: (A) an in-house AI agent built on LangChain and LangGraph, where the company’s engineering team (or a product studio) designs the orchestration graph, manages the model calls, and owns the integration code; and (B) a managed workflow-orchestration service delivered as a fixed-scope pilot, where a vendor such as Forfis scopes one back-office workflow, ships a human-in-the-loop pipeline in 8 weeks, and hands over a measured before/after baseline on cycle time and error rate. Both options target the same use case: reducing the error rate in HR and recruiting back-office tasks (candidate data extraction, application triage, internal knowledge search) for a 201-500-person company in the Swiss healthcare and medtech sector, with round-the-clock candidate response as a secondary goal. The comparison is not “build vs. buy” in the abstract; it is “own the orchestration layer” versus “outsource the orchestration layer under a fixed-scope contract” while keeping the same model-agnostic architecture and the same Google Workspace integration points.

    Seven Criteria for the Comparison

    We judge the two options against seven criteria that matter for a Swiss healthcare company running isolated pilots:

    • Time to first measurable result — weeks from kickoff to a working pipeline with a logged baseline.
    • Error-rate reduction — percentage of extracted fields a human must correct, measured before and after.
    • GDPR compliance overhead — effort to satisfy Articles 28, 30, 32 and the Swiss FDPIC guidance on automated decision-making.
    • Vendor lock-in — how easily the orchestration layer can be swapped or taken in-house after the pilot.
    • Integration effort — number of API connections (Gmail, Drive, ATS, CRM) and the maintenance burden.
    • Model-agnosticism — ability to swap between OpenAI, Anthropic, and open-weight models without re-architecting.
    • Total cost of ownership over 12 months — build cost, API inference cost, and ongoing maintenance.

    Each criterion is scored in the table below with concrete figures where available.

    Side-by-Side Comparison

    Criterion Option A: In-House LangGraph Agent Option B: Managed Fixed-Scope Pilot
    Time to first result 10-14 weeks (design, build, test, baseline) 8 weeks (fixed scope, pre-built integration templates)
    Error-rate reduction Depends on prompt engineering; typically 8-15% residual after 3 iterations 4-6% residual at pilot close-out, with logged human corrections
    GDPR compliance overhead Internal legal + engineering must map data flows, sign DPA, document Article 30 records Vendor provides DPA, data-flow map, and Article 30 log as pilot deliverables
    Vendor lock-in None — code is owned; LangGraph is open-source Low — orchestration graph is documented; model calls are API-based, not proprietary
    Integration effort 3-5 engineer-weeks for Gmail, Drive, ATS, CRM OAuth + API wiring Included in pilot scope; vendor maintains integration during the 8 weeks
    Model-agnosticism Full — swap any OpenAI/Anthropic/open-weight model at the node level Full — same architecture; vendor configures the model endpoint per workflow
    12-month TCO ~CHF 180 000-250 000 (1 FTE engineer + API costs ~CHF 4 000/month) ~CHF 95 000-130 000 (pilot fee + managed operation ~CHF 3 500/month)

    The TCO figures assume a single workflow with two integration points and moderate inference volume (roughly 500 candidate applications per month).

    When the In-House Agent Wins

    Option A wins when the company already has a dedicated engineering team of at least two full-time developers who can maintain the LangGraph codebase, write integration tests, and iterate on prompts after the pilot. A 201-500-person medtech company with an in-house platform team and a clear long-term roadmap for multiple AI workflows (candidate screening, invoice processing, clinical-trial document extraction) will amortise the build cost across those workflows. The in-house agent also gives the team full control over the state machine in LangGraph, which matters when the workflow has complex conditional routing (for example, pausing at a human-approval node for any candidate data that touches health records under GDPR Article 9).

    Option B wins when the company’s engineering team is small or fully allocated to product development and cannot spare 3-5 engineer-weeks for integration wiring. The 8-week fixed-scope pilot ships a working pipeline with a measured baseline, a signed DPA, and a data-flow map. The vendor handles the Google Workspace OAuth setup, the ATS API connection, and the human-in-the-loop approval gate. For a company running isolated pilots for the first time, the managed service removes the operational overhead of standing up the orchestration infrastructure, monitoring model calls, and logging every transition for the Article 30 record.

    Recommendation for the Swiss Healthcare Scenario

    Option B is the better fit for the stated scenario. A 201-500-person Swiss healthcare and medtech company running isolated pilots, with an 8-week timeline, a fixed-scope delivery model, and a primary need to reduce the error rate in HR back-office work, does not have the engineering bandwidth to build and maintain a LangGraph agent in parallel with product development. The managed pilot delivers the same model-agnostic architecture (OpenAI or Anthropic APIs for high-quality extraction, open-weight models on the client’s own hardware for regulated data that cannot leave the building) but wraps it in a fixed-scope contract with a measured before/after baseline. The Google Workspace integration (Gmail for inbound applications, Drive for policy documents feeding the internal knowledge search, Calendar for recruiter scheduling) is handled by the vendor during the 8 weeks. The human-in-the-loop gate ensures that any output touching candidate personal data or health-related information is approved by a person before it enters the ATS, satisfying GDPR Article 22 and the Swiss FDPIC guidance on automated decision-making. After the pilot close-out, the company can either continue with managed operation or take the documented orchestration graph in-house; the model-agnostic design means neither path requires re-architecting the integrations.

  • Document Extraction Pilot for E-Commerce Operations in Austria

    The Operational Bottleneck: Manual Order and Shipment Data Entry

    E-commerce and retail operations teams in Austria face a persistent bottleneck: order and shipment status updates from suppliers arrive in inconsistent formats—PDFs, scanned images, email attachments, and portal exports. Manual extraction and data entry into SAP or Microsoft Dynamics consumes 30-45 minutes per batch, with error rates averaging 2-4% that cascade into delayed customer notifications and reconciliation headaches.

    A fixed-scope pilot addresses this by automating one specific workflow within a four-week window. The engagement starts with a process audit that maps your current document flow, measures baseline cycle time and error rate, and identifies the highest-ROI extraction targets. From there, the team builds a document extraction pipeline using the OpenAI API for its strong performance on varied layouts, integrates it with your existing ERP via native APIs, and validates results against your baseline metrics.

    The deliverable is not a new system but a faster, more accurate version of the workflow you already run. Senior operations staff move from data entry to exception handling and supplier relationship management, while the AI layer handles the repetitive extraction and mapping work.

    Four-Week Pilot Structure: From Audit to Validated Pipeline

    The four-week timeline follows a structured sequence. Week one covers the process audit: the team reviews 50-100 sample documents from your supplier base, maps data fields to your ERP schema, and establishes the baseline metrics—current cycle time per batch, error rate, and staff hours consumed. This phase also confirms compliance requirements under the EU AI Act, including transparency logging and human oversight protocols for data that affects financial records.

    Weeks two and three handle model configuration and integration. The OpenAI API is tuned for your specific document types, with prompt engineering and post-processing rules to handle edge cases like merged invoices or multi-page shipments. The extraction pipeline connects to SAP or Microsoft Dynamics through their standard APIs, writing validated data directly to the relevant tables. Human-in-the-loop review queues are configured so that low-confidence extractions route to staff for approval before ERP sync.

    Week four focuses on validation and handover. The team processes a full week’s worth of live documents, compares results against the baseline, and documents the error rate, cycle time improvement, and any remaining edge cases. The handover package includes runbooks, model version records, and escalation procedures for ongoing managed operation.

    Model-Agnostic Architecture: OpenAI API and Open-Weight Options

    The architecture is deliberately model-agnostic, but the OpenAI API serves as the default for quality-critical extraction tasks. Its strength lies in handling varied document formats—scanned PDFs with mixed layouts, email attachments with inconsistent headers, and portal exports with variable column structures—without requiring custom OCR preprocessing for each format.

    For regulated data that cannot leave the building, the same pipeline runs on open-weight models deployed on your own hardware. This configuration maintains the same integration points and human-in-the-loop workflows while ensuring data sovereignty. The trade-off is higher initial setup effort and potentially lower accuracy on edge cases, which the human review queue compensates for.

    The pipeline plugs into your existing SAP or Microsoft Dynamics ERP through their standard APIs rather than replacing them. Extracted data maps to your existing data structures: order numbers to sales order tables, shipment dates to delivery schedule lines, status codes to your internal workflow states. No ERP migration or reconfiguration is required. The AI layer sits alongside your current systems, handling the extraction and mapping work while your ERP continues to manage the downstream business logic.

    EU AI Act Compliance: Transparency and Human Oversight

    The EU AI Act classifies document extraction systems as limited-risk AI, requiring transparency about AI involvement and human oversight for decisions that affect financial records or customer commitments. For e-commerce operations in Austria, this means the system must log its actions, maintain records of model versions and training data, and allow human review before extracted data syncs to the ERP.

    The pilot ships with compliance documentation built in: action logs showing which documents were processed, confidence scores for each extraction, and a review trail for any human approvals. Model version records track which API version or open-weight model was used for each batch, supporting audit requirements. The human-in-the-loop workflow ensures that anything touching money, health data, or contracts requires explicit staff approval before ERP sync.

    For a 501-2000 employee company, this compliance layer adds minimal overhead to the four-week timeline. The documentation and logging are configured during the integration phase, and the review queue is part of the standard human-in-the-loop setup. The result is a system that meets EU AI Act requirements without requiring a separate compliance project or legal review cycle.

    Measuring Success: Cycle Time, Error Rate, and Staff Hours

    The pilot’s success is measured against the baseline established in week one. Typical targets for order and shipment status extraction include reducing cycle time from 30-45 minutes per batch to under 10 minutes, cutting error rates from 2-4% to under 0.5%, and freeing 60-80% of the staff hours previously consumed by manual data entry.

    The before/after comparison uses the same document samples processed through both the manual and AI-assisted workflows. Cycle time measures the elapsed time from document receipt to ERP sync. Error rate counts the number of fields requiring correction after initial extraction, divided by total fields processed. Staff hours are tracked through time-stamped review queues, showing how much time staff spend on exception handling versus routine data entry.

    The handover package includes a validation report with these metrics, a runbook for daily operations, and escalation procedures for edge cases. The managed operation phase continues with monthly performance reviews, model updates as supplier document formats change, and support for new document types as your supplier base evolves. The goal is not a one-time automation but a continuously improving AI-native operations layer that scales with your business.

  • UK Fintech Cuts Support Ticket Cost 30-40% with AI Document Extraction Pilot

    Background: A 25-Person UK Fintech at the Pilot Stage

    This case study is a composite based on patterns observed in the field. We do not fake named customers. The details reflect real engagement structures, technical constraints, and outcome ranges we have seen across multiple fintech and payments clients in Tier-1 markets.

    The company in question is a 25-person fintech operating in the UK, focused on payment processing for small and medium businesses. They run a lean sales and support team that handles inbound leads, processes support tickets, and manages customer relationships through a CRM. Their stack includes a commercial CRM, a helpdesk platform, and a custom payment processing backend. The team is at the ‘Running Isolated Pilots’ stage of AI maturity, meaning they have experimented with AI tools but have not yet integrated them into core workflows. They recognize the value of AI but lack the process to implement it systematically.

    Challenge: Multilingual Support and Lead Qualification Under Pressure

    The company faced three operational pressures simultaneously. First, their support team was handling tickets in English, Spanish, and French, but they only had two multilingual staff members. This created bottlenecks and increased cost per support ticket. Second, their sales team was manually qualifying inbound leads from web forms and email, a process that took 4-6 hours per lead and delayed response times. Third, they were preparing for an ISO 27001 audit and needed to demonstrate that any new systems would meet their compliance requirements.

    The deadline was tight: they needed to show measurable improvements within 8 weeks to justify the investment to their board. The headcount constraint was real, as they could not hire additional multilingual staff without significantly increasing their operating costs. The compliance requirement added another layer of complexity, as any AI system they deployed would need to handle sensitive financial data and customer contracts with appropriate safeguards.

    Approach: AI Automation Audit and Document Extraction Pipeline

    The team engaged Forfis to run an AI automation audit, a structured process that maps existing workflows, identifies the highest-impact automation opportunities, and designs a fixed-scope pilot. The audit took two weeks and produced a prioritized list of workflows to automate. The top two were document extraction for inbound lead forms and support tickets, and multilingual classification for lead qualification.

    The technical approach used the OpenAI API for its strong multilingual capabilities and accuracy in document extraction. The team built custom REST API endpoints and webhooks to integrate with their existing CRM and support systems. The architecture was deliberately model-agnostic, allowing them to swap in open-weight models later if data residency requirements changed. Human-in-the-loop approval was built in for any data touching financial records or customer contracts. The system never stored raw documents longer than 72 hours, and all processing occurred within the UK data residency boundary.

    Outcome: Measurable Improvements in 8 Weeks

    The 8-week timeline included two weeks for the process audit and workflow mapping, three weeks for building and testing the document extraction pipeline, and three weeks for integration, pilot testing, and baseline measurement. The team shipped a measured before/after comparison on cycle time and error rate.

    The results were concrete. Cost per support ticket dropped by 30-40%, as the automated extraction reduced manual data entry time. Lead qualification speed improved by 25-35%, as the system classified and routed leads in minutes rather than hours. Manual data entry time decreased by 15-20%, freeing the support team to focus on complex issues. The error rate in data extraction was 2-3%, well within the acceptable range for their use case. These metrics were tracked over a four-week pilot period with human oversight on all sensitive data.

    Lessons for Similar Fintech Teams

    • Start with a fixed-scope pilot, not full automation. The team focused on one workflow (document extraction) rather than attempting to automate all support and sales processes. This reduced risk and built confidence for rollout.
    • Maintain human-in-the-loop approval for sensitive data. Any extracted data touching financial records or customer contracts required manual review before entering the CRM. This maintained ISO 27001 compliance and built trust with the team.
    • Build the architecture to be model-agnostic. The team used the OpenAI API for its strong multilingual capabilities but designed the system to swap in open-weight models if data residency requirements changed. This future-proofed the investment.
    • Measure baseline metrics before and after the pilot. The team tracked cycle time, error rate, cost per ticket, and lead qualification speed. These concrete numbers justified the investment and provided a clear path to rollout.
    • Integrate with existing systems, not replace them. The custom REST API and webhooks kept the integration lightweight and avoided the cost and risk of replacing the CRM and helpdesk.
  • Forfis AI Automation for Lead Qualification in German Professional Services

    Process Audit and Fixed-Scope Pilot

    Professional services firms in Germany with 201-500 employees face a specific bottleneck: manual data entry and slow lead response erode margins. The process audit identifies which workflows are worth automating, typically lead qualification and document extraction. The pilot targets one workflow, not enterprise-wide transformation, keeping scope fixed and results measurable. The architecture plugs into existing CRMs, ERPs, and helpdesks through their native APIs rather than replacing them. Forfis uses OpenAI and Anthropic APIs where quality matters and open-weight models on the client’s own hardware where data cannot leave the building. The human-in-the-loop default means the model drafts or classifies, but a person approves anything touching contracts or financial commitments. Every pilot ships with a measured before/after baseline on cycle time and error rate to prove value before rollout.

    Conversational Agent and Document Extraction Pipeline

    The document extraction pipeline processes inbound PDFs, spreadsheets, and email attachments to pull structured data into your CRM. The conversational agent handles the first touch: it answers FAQs, captures intent, and routes tickets. The agent uses the extracted data to personalize follow-ups and qualify leads based on predefined criteria. For lead qualification, the agent auto-responds to standard inquiries but flags complex or high-value leads for human review. The architecture is deliberately model-agnostic, using OpenAI and Anthropic APIs where quality matters and open-weight models on the client’s own hardware where regulated data cannot leave the building. The integration layer abstracts the model choice, so you can switch providers without rebuilding the pipeline. The human-in-the-loop approval process ensures that anything touching money, health data, or contracts requires human sign-off.

    8-Week Integration Sprint Timeline

    The integration sprint runs in parallel with your existing operations. Week 1-2 covers process audit and baseline measurement. Week 3-5 builds the pilot on one workflow, typically lead qualification or document processing. Week 6-7 tests with real data and human-in-the-loop approval. Week 8 documents results and plans rollout. No systems are replaced during the sprint. The AI layer connects to Notion or Confluence through their APIs to retrieve company documentation, pricing sheets, and service descriptions. This allows the conversational agent to answer questions with accurate, up-to-date information from your own knowledge base. The retrieval-augmented approach ensures responses reflect your current offerings, not generic training data. The pilot ships with a measured baseline on cycle time and error rate to prove value before rollout.

    Measuring Success: Cycle Time and Error Rate Baselines

    The pilot targets one workflow to keep scope fixed and results measurable. Success means the AI layer reduces manual data entry by a measurable percentage and improves response time. For lead qualification, the target is typically a 30-50% reduction in time-to-first-response and a 20-40% improvement in lead accuracy. For document extraction, the target is a 40-60% reduction in processing time and a 15-30% improvement in data accuracy. The pilot ships with a measured before/after baseline on cycle time and error rate to verify the human-in-the-loop process works as intended. The architecture plugs into existing CRMs, ERPs, and helpdesks through their native APIs rather than replacing them. The model-agnostic design means you can choose the model based on your data sensitivity and quality requirements without rebuilding the pipeline.

    Scaling Across Departments After the Pilot

    The pilot focuses on one workflow to keep scope fixed and results measurable. Rollout to additional departments happens after the pilot proves value, typically in 4-6 week increments. Each new department gets its own baseline measurement and human-in-the-loop approval process. Scaling across departments is a phased process, not a big-bang deployment. The architecture is deliberately model-agnostic, using OpenAI and Anthropic APIs where quality matters and open-weight models on the client’s own hardware where regulated data cannot leave the building. The integration layer abstracts the model choice, so you can switch providers without rebuilding the pipeline. The human-in-the-loop default means the model drafts or classifies, but a person approves anything touching contracts or financial commitments. Every rollout includes a measured before/after baseline on cycle time and error rate to prove value before expanding to the next department.

  • AI Process Audit vs. Cost-per-Ticket Reduction: A Fintech Comparison

    What Is Being Compared

    The two options under evaluation are not competing products but competing entry points into the same AI automation program. Option A, the AI process audit and roadmap, is a diagnostic engagement: Forfis maps the company’s existing workflows, measures cycle time and error rate on each, scores them by volume and data sensitivity, and delivers a 12-month automation roadmap with a fixed-scope pilot on the highest-ROI workflow. Option B, lower cost per support ticket, is an outcome-oriented engagement: the client specifies a target reduction in cost per ticket (e.g., 40% over two quarters), and Forfis designs the AI layer—triage, first-response, predictive scoring—directly against that KPI. Both engagements use the same delivery stack: n8n orchestration, model-agnostic LLM integration, Google Workspace connectors, and human-in-the-loop approval gates. The difference is where the engagement starts: from the process map or from the P&L line.

    Criteria for Judgment

    The comparison is judged against eight criteria that matter to a 501-2000 employee fintech operating under PCI DSS in Germany:

    • Time to first measurable result — weeks from kickoff to a quantified before/after baseline
    • PCI DSS compliance surface — how much cardholder data touches the AI layer
    • n8n orchestration depth — how many workflow nodes, conditional branches, and API calls the solution requires
    • Predictive scoring accuracy — AUC or F1 on the lead-qualification model at pilot exit
    • Multilingual coverage — number of languages supported in the first release
    • Google Workspace integration — email, calendar, and document access from the AI agent
    • Cost per support ticket — measured reduction against the pre-pilot baseline
    • Managed AI Operations scope — what Forfis operates post-go-live versus what the client’s team owns

    Comparison Table

    Criterion Option A: AI Process Audit and Roadmap Option B: Lower Cost per Support Ticket
    Time to first measurable result 4 weeks (pilot go-live on one workflow) 4 weeks (pilot go-live on support triage)
    PCI DSS compliance surface Low — audit phase touches no CHDE; pilot workflow selected to avoid CHDE Medium — support tickets may reference transaction IDs; n8n workflow masks CHDE before LLM call
    n8n orchestration depth 15-25 nodes (audit scoring, routing, baseline measurement) 25-40 nodes (ticket classification, first-response drafting, escalation, CRM update)
    Predictive scoring accuracy N/A in audit phase; scored in roadmap for future workflows F1 ≥ 0.82 on lead-qualification subset at pilot exit
    Multilingual coverage 1 language (English) in pilot; roadmap adds 2-3 languages in months 2-3 2 languages (English, German) in pilot; additional languages in month 2
    Google Workspace integration Read-only access to email and calendar for audit context Read/write access for first-response drafting and ticket status updates
    Cost per support ticket Not the primary KPI; measured as secondary metric Primary KPI; target 35-50% reduction by month 3
    Managed AI Operations scope Forfis operates n8n workflows, model monitoring, and roadmap execution Forfis operates n8n workflows, model monitoring, ticket KPI reporting, and escalation handling

    Scenario-by-Scenario Verdict

    Option A wins when the company has no clear starting point. A fintech with 501-2000 employees often runs 15-30 back-office and customer-facing workflows, and the leadership team cannot tell which one will yield the fastest ROI. The audit resolves that ambiguity: Forfis measures cycle time and error rate on each candidate, scores them against volume and data sensitivity, and delivers a ranked roadmap. The 4-week pilot then targets the top-ranked workflow—often lead qualification in a payments company, because it has high volume, measurable conversion data, and no direct CHDE exposure. The roadmap gives the CFO a 12-month view of cumulative savings, which is what unblocks budget for subsequent phases.

    Option B wins when the company already knows the problem. If the support desk is handling 3,000-5,000 tickets per month at an average cost of EUR 12-18 per ticket, and the VP of Customer Experience has a board-level target to cut that by 40%, the audit phase is redundant. The engagement starts directly on the support workflow: n8n classifies each incoming ticket, the LLM drafts a first response, a human approves anything touching a refund or a contract clause, and the system logs cycle time and error rate against the pre-pilot baseline. The 4-week timeline is tighter because the scope is fixed from day one.

    Recommendation

    For a German fintech with 501-2000 employees operating under PCI DSS, the recommendation depends on one question: does the leadership team have a named KPI with a target number? If yes—“cut cost per support ticket by 40% by Q3”—start with Option B. The 4-week pilot on support triage delivers a measurable baseline, the n8n workflow is scoped to the ticket lifecycle, and the PCI DSS data-flow review is contained to the support system. Multilingual coverage (English and German) ships in the pilot; additional EU languages follow in month 2.

    If the answer is no—if the company knows AI can help but cannot say where—start with Option A. The audit identifies the highest-ROI workflow, the roadmap sequences the next three, and the 4-week pilot proves the delivery model. For a company in this size range, the audit typically surfaces lead qualification as the first pilot because it sits at the intersection of marketing and revenue, touches no CHDE, and has a clean before/after metric (conversion rate, time-to-first-response). The predictive scoring model, built on historical lead data, reaches F1 ≥ 0.82 by pilot exit and feeds the n8n routing logic that sends high-score leads to human SDRs within 2 hours.

  • UK B2B SaaS Firm Cuts Invoice Cycle Time 74% With On-Premise AI Pilot

    Background: A 30-Person B2B SaaS Firm in Manchester

    This case study is a composite drawn from patterns observed across multiple engagements. No named customer is represented. The firm described here is a 30-person B2B SaaS company based in Manchester, selling a project-management platform to mid-market clients across the UK and Ireland. The operations team of four handles supplier invoices, delivery notes, and credit notes for a mix of cloud hosting, office supplies, and professional services vendors. The existing stack is a standard ERP (Xero for accounting, a lightweight project-management tool for internal tracking) and Slack as the primary communication channel. No AI system is in production anywhere in the company. The trigger for change is not a technology initiative but a headcount constraint: the operations lead has been absorbing invoice processing work that was previously split across two part-time staff, and the founder has set a deadline to reduce the manual workload before the next hiring cycle in Q3.

    Challenge: Four-Day Cycle Time and a GDPR Gap

    The operations lead processes roughly 180 supplier invoices per month, each requiring manual data entry into Xero: vendor name, line items, tax codes, and total amount. The median cycle time from invoice receipt to payment approval is four business days, with a long tail of invoices taking nine to twelve days when the operations lead is pulled into client escalations. The error rate on manual data entry is 18 percent, measured over a two-week sample in the audit phase. Each error triggers a correction cycle that adds 20 to 35 minutes of senior staff time. The compliance pressure is GDPR: the invoices contain personal data (vendor contact names and email addresses), and the firm’s data protection officer has flagged that the current manual process, which involves forwarding PDFs between personal email accounts and the operations lead’s inbox, does not meet the Article 5(1)(f) integrity and confidentiality requirement. The deadline is eight weeks: the founder wants a working pilot before the Q3 hiring decision, and the data protection officer wants a documented DPIA before any new system touches the invoice data.

    Approach: Two-Week Audit, Fixed-Scope Pilot, On-Premise Inference

    The engagement starts with a two-week AI automation audit. The team maps every document that enters the operations workflow, measures the current cycle time and error rate, and scores each workflow on volume, error cost, and automation feasibility. Invoice processing wins the composite score: 180 documents per month, a 18 percent error rate with a 20-to-35-minute correction cost per error, and a document format that maps cleanly to a structured extraction task. The pilot scope is fixed: extract vendor name, line items, tax codes, and total amount from PDF invoices, write the data to Xero via the API, and route flagged fields to the operations lead in Slack for approval. The architecture is model-agnostic: the orchestration service routes inference to an on-premise vLLM endpoint running a 7B-parameter open-weight model, because the GDPR review confirms that the invoice data cannot be sent to a cloud API. The Slack integration is built with the Slack Bolt framework, posting flagged items to a dedicated channel with approve and reject buttons. The human-in-the-loop gate is hard-coded: any field with a confidence score below 0.92 is flagged for human review.

    Outcome: 74 Percent Cycle-Time Reduction in Six Weeks

    The pilot runs for six weeks after the audit, with a two-week shadow period at the end where the AI drafts and the operations lead approves every output. The before/after baseline is measured over the final two weeks of the shadow run. The median cycle time drops from 4.2 days to 1.1 days, a 74 percent reduction. The manual correction rate falls from 18 percent to 4 percent. The operations lead reviews 22 flagged items per day in week one, dropping to 8 per day by week six as the model’s confidence improves on the firm’s specific vendor set. The senior operations lead, who had been spending roughly 14 hours per week on invoice processing, reports spending 3 hours per week on the approval queue and 2 hours per week on exception handling. The GDPR DPIA is completed in week three, documenting the data flows, the retention policy (invoices retained for seven years per UK tax law, extracted data retained for 12 months), and the human-in-the-loop approval gate. The on-premise hardware is a single workstation with an NVIDIA L40S 48 GB GPU, provisioned in week one and running the vLLM inference server for the duration of the pilot.

    Lessons for Similar Teams

    • The audit is the product, not the pilot. The two-week process audit produced a one-page baseline report that the client retained for internal reporting and the GDPR accountability record. The pilot was the validation, but the audit was the deliverable that justified the investment. Teams that skip the audit and jump straight to a pilot often discover mid-engagement that the workflow they chose is not the highest-impact one. – On-premise hardware is a procurement decision, not a technical one. The L40S workstation was ordered in week one, before the audit was complete. The lead time for GPU hardware in the UK is four to six weeks. Teams that order the hardware after the audit is done lose two to three weeks of the pilot timeline. – The Slack integration is the adoption lever. The operations lead approved 22 items per day in week one without any training, because the interface was the tool she already used. A separate dashboard would have added friction and likely reduced the approval rate below the threshold needed for the baseline comparison. – The confidence threshold is a tuning parameter, not a fixed constant. The 0.92 threshold for monetary fields was set in week one and adjusted to 0.95 in week four after the model’s performance on the firm’s specific vendor set improved. Teams that treat the threshold as a fixed constant either over-flag (wasting senior time) or under-flag (letting errors through). – The GDPR DPIA is a two-week task, not a one-day checkbox. The data protection officer spent three hours in week two reviewing the data flow diagram and two hours in week three reviewing the retention policy. The DPIA was completed in week three, not week one, because the model’s training data provenance had to be documented before the review could be signed off.
  • 8-Week RAG Candidate Screening Pilot for a German E-commerce Team

    The problem: manual screening and reporting eat your HR team’s week

    You run an e-commerce or retail operation in Germany with 11 to 50 employees. Your HR and recruiting team spends 6 to 10 hours per week manually screening CVs, extracting skills and experience into a spreadsheet, and matching candidates against job postings. The monthly reporting cycle compounds the problem: you pull data from the ATS, reconcile it with the spreadsheet, and format a report for leadership, all by hand. The goal is not to replace the recruiter but to cut the manual back-office work around screening and reporting, so the team spends time on interviews and hiring decisions instead of data entry. The constraint is that candidate data is personal data under GDPR, and your ISO 27001 certification requires documented access controls and audit trails. The pilot must prove a measurable reduction in cycle time and error rate within 8 weeks, using the OpenAI API for the model layer and a custom REST API with webhooks to connect to your existing ATS and reporting tools.

    Prerequisites before week one

    Before the pilot starts, confirm the following are in place:

    • A working ATS or candidate log. Even a structured spreadsheet with columns for name, email, skills, experience, and job applied to qualifies. The pipeline needs a defined schema to write results back to.
    • A set of 10 to 30 active job postings with written competency requirements. These become the RAG index source. If your job descriptions are vague, the model will match vaguely.
    • A named data owner who can approve the data-processing agreement for the OpenAI API and sign off on the ISO 27001 security annex.
    • A 200-sample gold set of past CVs with manually verified extraction fields. This is your error-rate baseline. Without it, you cannot measure whether the pipeline is accurate.
    • API access to your ATS or reporting tool, or a willingness to expose a minimal REST endpoint. The pilot integrates through custom REST API and webhooks, not by replacing your existing system.
    • A point of contact who can approve scope changes within 48 hours. Fixed-scope means the SOW is locked after week one; slow approvals stall the timeline.

    Step 1: Run the process audit and capture the baseline

    Spend the first five business days mapping the current workflow. Have the HR team process a sample batch of 50 CVs manually and time each step: receipt, initial read, field extraction, matching against the job posting, and entry into the log. Record the cycle time in minutes per CV and the error rate by having a second person verify the extracted fields. This baseline is the denominator for every metric in the week-8 report. Simultaneously, inventory the document types you receive: PDFs, DOCX, scanned images, and email attachments. Note which fields vary by job type. The audit output is a one-page process map with timestamps and a list of the top five error categories. This document becomes the scope anchor for the pilot SOW.

    Step 2: Build the document extraction pipeline

    Build the extraction pipeline to parse incoming CVs into structured JSON. Use a document parser such as Apache Tika or a cloud OCR service for scanned PDFs, then feed the text to the OpenAI API with a system prompt that specifies the target schema: name, email, phone, skills (array), years_experience (number), education (array of objects), and job_titles (array). The prompt should include two or three few-shot examples from your gold set to anchor the output format. Log every API call with the input hash, the model version, the response, and a timestamp. Store the structured output in a staging table. The pipeline should handle a batch of 20 CVs in under 90 seconds at the OpenAI gpt-4o token rate, which is roughly 120 tokens per CV for a typical one-page document. If a CV fails to parse, flag it for manual review rather than guessing.

    Step 3: Build the RAG index over your job postings

    Index your job postings, competency matrices, and past hiring decisions into a vector store. Use a chunking strategy that keeps each job requirement as a separate chunk so the RAG retrieval can cite specific criteria. Embed the chunks with a model such as text-embedding-3-small from OpenAI and store them in a vector database like Weaviate or Qdrant running on your own infrastructure, since the job-posting data may contain internal compensation bands or hiring criteria you do not want in a third-party vector service. The RAG query flow is: take the extracted candidate profile, generate a query string, retrieve the top 5 most relevant job-requirement chunks, and pass them to the OpenAI API with a prompt that asks the model to score the match from 0 to 100 and cite which specific requirements were met or missed. The output is a JSON object with the score, the cited requirements, and a one-paragraph rationale.

    Step 4: Wire the REST API and webhooks to your ATS

    Expose three REST endpoints: POST /documents to upload a CV, GET /jobs/{id} to retrieve a job posting’s indexed criteria, and POST /results to submit the classification back to your ATS. Configure webhooks so that when the pipeline finishes processing a batch, it fires a batch.completed event to your integration layer with a payload containing the correlation ID, the list of candidate references, the average confidence score, and a link to the full output. Your ATS or integration layer acknowledges with a 200 response within 5 seconds. If it does not, the pipeline retries with exponential backoff: 10 seconds, 30 seconds, 90 seconds. After three failed retries, the record is flagged in the review queue with a webhook_failed status. The human-in-the-loop step sits here: a recruiter sees the model’s score, the cited requirements, and the raw CV side-by-side, and clicks approve or reject. Every approval or rejection is logged with the recruiter’s user ID and timestamp for the ISO 27001 audit trail.

    Step 5: Run the pilot with human-in-the-loop review

    Run the pipeline on a live batch of 50 to 100 CVs over two weeks. The recruiter reviews every classification, and you log each correction: which field was wrong, what the model said, and what the correct value was. At the end of the run, compute the error rate against the gold set and compare it to the baseline from step 1. If the error rate is above 5 percent, identify the top three error categories and adjust the extraction prompt or the RAG retrieval parameters. Common fixes: tighten the few-shot examples, add a negative constraint to the prompt (“do not infer skills that are not explicitly stated”), or increase the number of retrieved chunks from 5 to 8. Re-run the batch after each adjustment. The goal is to bring the error rate under 5 percent and the cycle time under 30 seconds per CV before the week-8 report. Document every prompt change and its effect in a change log.

  • AI Ticket Triage Glossary for Swiss Medtech: 12 Terms from Pilot to Rollout

    A-D: Core Workflow Terms

    The following terms are defined in the context of a 51-200 employee Swiss medtech company deploying AI-assisted ticket triage and data enrichment for the first time. The company has no AI in production, operates under Swiss FADP and EU AI Act obligations, and runs open-weight models on-premise to keep patient data within the building. Each entry includes a definition and a contextual example drawn from this scenario.

    Ticket Triage and Routing is the classification and assignment of incoming support tickets by urgency, topic, and required expertise. In a medtech firm, this distinguishes a firmware bug report from a patient safety alert. An AI system classifies each ticket in under 30 seconds; a human reviews any ticket flagged as high-risk before it reaches a clinical team.

    Document and Data Extraction Pipelines are automated workflows that pull structured fields from unstructured sources like PDFs and emails. For this company, the pipeline extracts device serial numbers and error codes from incoming tickets and writes them to the CRM via REST API, replacing 2-4 hours of daily manual re-entry.

    E-M: Architecture and Integration Terms

    These terms describe the technical architecture and integration approach for a compliance-constrained deployment.

    Open-Weight Models On-Premise refers to running publicly available model weights (Llama 3, Mistral, Falcon) on the company’s own hardware. For a Swiss medtech firm, this ensures patient data never leaves the building, satisfying FADP and EU AI Act data residency requirements. The trade-off is that open-weight models require more tuning than proprietary APIs but perform reliably for structured classification and extraction tasks.

    Custom REST API and Webhooks are the integration layer connecting the AI system to existing CRMs, ERPs, and helpdesks. When a new ticket arrives, a webhook fires; the AI classifies it; the result is pushed back via REST API. This preserves existing user interfaces and reduces change management friction for a team of 51-200 employees who already know their tools.

    Data Enrichment and Cleanup is the process of augmenting raw ticket data with CRM and ERP records (device serial, firmware version, prior support history) and normalizing inconsistent formats. This step ensures the AI and downstream processes work with clean, complete data rather than the messy input that manual entry produces.

    N-R: Compliance and Delivery Terms

    These terms cover the regulatory and delivery framework governing the rollout.

    EU AI Act is the European Union’s regulation of AI systems, classifying those affecting health, safety, or legal rights as high-risk. Article 14 mandates human oversight for high-risk systems. For a Swiss medtech firm serving EU customers, the Act applies extraterritorially, requiring documented risk assessments, transparency logs, and human sign-off for any routing decision involving patient safety.

    Fixed-Scope Pilot is a time-boxed engagement (4 weeks in this scenario) with predefined deliverables, success metrics, and a hard stop. The scope is locked before work begins: the specific workflow, data sources, integration points, and baseline measurements. For a company with no prior AI deployment, this model limits financial risk and provides a measurable before/after comparison on cycle time and error rate.

    Process Audit is the structured review of existing workflows to identify which tasks are repetitive, error-prone, and suitable for automation. It maps who does what, how long each step takes, and where errors occur. For a firm with no AI in production, this audit prevents the common mistake of automating a broken process and ensures the pilot targets the workflow with the highest ROI.

    S-Z: Operational and Organizational Terms

    These final terms describe the operational and organizational context of the deployment.

    Human-in-the-Loop (HITL) is a design pattern where a human reviews and approves AI-generated outputs before they take effect. For a medtech company, any ticket routed to a clinical team, any data entry involving patient records, and any response touching a contract requires human sign-off. The AI drafts, classifies, or extracts; the human validates. This satisfies EU AI Act Article 14 and builds organizational trust during the transition from manual to automated workflows.

    Compliance-Safe AI Rollout is a phased deployment strategy ensuring regulatory requirements are met at every stage. It starts with a risk assessment, proceeds to a fixed-scope pilot with human oversight, and scales only after the pilot demonstrates measurable improvements without compliance breaches. For a Swiss medtech firm, this means documenting every AI decision, maintaining audit logs, and ensuring the on-premise architecture prevents data exfiltration.

    No AI in Production Yet means the company has no deployed AI systems handling live business processes. The pilot must therefore include foundational setup: model deployment, API integration, baseline measurement, and staff training, all within the 4-week timeline.