Author: Forfis

  • UK SaaS Team Cuts Candidate Screening from 18 Days to 6 in a Four-Week AI Pilot

    Background: A 30-Person UK SaaS Team with a Screening Bottleneck

    This case study is a composite based on patterns observed across multiple engagements. We do not name real customers. The company, the metrics, and the timeline are drawn from a recurring profile: a 30-person B2B SaaS firm in the UK, mid-growth stage, running on a standard stack of Notion for documentation, a CRM for pipeline, and a helpdesk for support. The team had no dedicated AI function. The founder had read about LLMs and wanted to test whether one process could be automated without a six-month build. The engagement ran for four weeks, end to end, from process audit to measured baseline.

    The Challenge: 18-Day Screening Cycle and a Hiring Deadline

    The team was hiring for two roles simultaneously: a senior engineer and a customer success manager. The screening process was manual. A recruiter read each CV, wrote a summary in Notion, and flagged the candidate for the hiring manager. The average cycle time from application to first screening decision was 18 days. The error rate was not measured, but the hiring manager reported that roughly one in five candidates who passed screening were later found to be a poor fit. The pressure was operational: the founder needed to close both roles before the next funding round, and the manual process was the bottleneck. There was no compliance constraint, but the team wanted a clean, auditable trail of who approved each screening decision.

    Approach: Audit, Build, and a Human-in-the-Loop Gate

    The engagement started with a three-day process audit. The dedicated AI team mapped the screening workflow step by step, identified the two highest-impact automation points (CV extraction and screening summary), and selected candidate screening as the single pilot process. The build used Anthropic Claude API for the extraction and classification. The integration was read-write against Notion: the AI read the job description and the CV, wrote the screening summary back to the same Notion page, and tagged the candidate with a classification label. The human-in-the-loop step was a simple approve/edit/reject button on the Notion page. The multilingual coverage was built in from day one: the model handled CVs in English, French, and German without a separate translation step. The build took nine days. The remaining time was spent on the baseline measurement and the rollout to the two open roles.

    Outcome: 18 Days to 6 Days, 20% to 8% Error Rate

    The measured baseline showed a cycle time reduction from 18 days to 6 days for the screening step. The error rate, measured as the percentage of candidates who passed screening but were later rejected at interview, dropped from 20% to 8%. The human-in-the-loop step added 3 minutes per candidate, but the total time per candidate fell from 22 minutes to 9 minutes. The team screened 47 candidates in the four-week window, compared to 19 in the previous four weeks. The founder reported that the hiring manager could now review all screening decisions in a single 30-minute session per day, instead of spreading them across the week. The multilingual coverage meant the team could accept applications from candidates in France and Germany without a separate translation step, which the founder estimated saved roughly 4 hours per week.

    Lessons for Similar Teams

    • Start with one process, not a platform. The pilot succeeded because the scope was a single workflow with a clear input and output. Teams that try to automate three processes in four weeks end up with three half-built integrations and no clean baseline. – Measure the baseline before you build. The 18-day cycle time and 20% error rate were recorded in the first week. Without that number, the outcome would have been anecdotal. The baseline is the most valuable deliverable in the pilot. – Human-in-the-loop is not a compromise; it is the product. The approve/edit/reject gate is what made the hiring manager trust the output. Remove it, and the team reverts to manual screening within two weeks. – Multilingual coverage is a feature, not a nice-to-have. For a UK team hiring in a European market, the ability to screen CVs in French and German without a translation step is a direct operational gain. Build it in from day one. – The integration is the moat, not the model. The AI layer plugs into Notion through its API. If the team later switches to Confluence, the integration work is a day, not a rebuild. The model is swappable; the integration is the asset.
  • OpenAI API vs On-Prem Models for a Swiss Fintech Pilot

    What Is Being Compared

    The two options under comparison are the OpenAI API as a hosted inference service and an open-weight model running on the client’s own hardware. The OpenAI API is a managed service where prompts are sent over HTTPS and completions are returned; the client does not manage the model weights or the inference infrastructure. The on-prem option uses a model such as Llama 3 or Mistral, deployed on the client’s servers or a private cloud, where the model weights are downloaded and the inference runs locally. Both options can serve the same two workflows: a retrieval-augmented knowledge assistant over Confluence or Notion, and a ticket triage and routing system for the helpdesk. The comparison is framed for a Swiss fintech with 501 to 2000 employees, operating under ISO 27001, with a two-week fixed-scope pilot as the delivery vehicle. The goal is to free senior staff from routine work in operations and supply chain, specifically by reducing manual back-office tasks and automating first-response triage.

    Criteria for the Comparison

    The evaluation uses seven criteria that matter to a Swiss fintech under ISO 27001. Latency is measured as the time from prompt submission to first token, which affects the user experience in a RAG assistant. Cost per unit is the total expense per ticket triaged or per document extracted, including API fees, compute, and human review time. Data residency is whether the data leaves the client’s network, which is a hard constraint for payment data under FINMA guidance. Compliance fit is how well the option aligns with ISO 27001 controls, particularly access control, logging, and data processing agreements. Integration effort is the number of API calls and configuration steps needed to connect to Confluence, Notion, and the helpdesk. Model quality is measured on a defined evaluation set of 200 tickets and 100 documents, scored by a human reviewer. Vendor lock-in is the cost and effort of switching to a different model or provider after the pilot. Each criterion is scored in the table below with concrete numbers where available.

    Comparison Table

    Criterion OpenAI API On-Prem Open-Weight Model
    Latency (first token) 180 to 400 ms over HTTPS 50 to 150 ms on local GPU
    Cost per ticket triaged 0.02 to 0.05 USD per ticket 0.005 to 0.02 USD per ticket after amortized hardware
    Data residency Data leaves client network to OpenAI infrastructure Data stays on client hardware
    ISO 27001 fit Requires DPA and data flow documentation Easier to document; no external data transfer
    Integration effort 3 to 5 API calls; standard HTTPS 8 to 12 steps; requires GPU provisioning and model loading
    Model quality (200-ticket eval) 92 percent accuracy on triage 85 to 88 percent accuracy on triage
    Vendor lock-in Low; prompt templates are portable Low; model weights are open, but inference stack is tied to hardware

    The latency difference is small for batch processing but noticeable in a live RAG assistant where the user is waiting for a response. The cost difference is significant at scale: for 10,000 tickets per month, the OpenAI API costs 200 to 500 USD, while the on-prem model costs 50 to 200 USD after the initial hardware investment. The data residency row is the deciding factor for a fintech handling payment data.

    When the OpenAI API Wins

    For ticket triage and routing, the OpenAI API wins on quality and speed of deployment. The 92 percent accuracy on the 200-ticket evaluation set means fewer misroutes, which directly reduces the time senior staff spend correcting errors. The 180 to 400 ms latency is acceptable for a triage system where the user is not waiting for a real-time response; the ticket is routed asynchronously. The integration effort is lower: three to five API calls to the helpdesk and the OpenAI endpoint, with no GPU provisioning. For a two-week pilot, this means the team can focus on the classification logic and the human-in-the-loop approval step rather than on infrastructure setup. The cost of 0.02 to 0.05 USD per ticket is negligible at the pilot scale of a few hundred tickets.

    When the On-Prem Model Wins

    For the retrieval-augmented knowledge assistant over Confluence or Notion, the on-prem model is the stronger choice when the indexed documents contain payment data, customer identifiers, or internal financial records. The data residency constraint is non-negotiable: FINMA guidance for Swiss fintechs requires that personal data and payment data be processed within the client’s control. The on-prem model keeps the embeddings and the prompts on the client’s hardware, so no data leaves the building. The 50 to 150 ms latency is faster than the OpenAI API, which improves the user experience in a live assistant. The 85 to 88 percent accuracy is lower than the OpenAI API, but for a RAG assistant the quality is more dependent on the retrieval step than on the model itself. The integration effort is higher, requiring GPU provisioning and model loading, but this is a one-time setup that pays off over the life of the assistant.

    Recommendation for the Swiss Fintech Pilot

    The recommendation is a hybrid architecture that uses the OpenAI API for ticket triage and the on-prem model for the RAG assistant. This split is driven by the data residency constraint: ticket data in a helpdesk is less sensitive than the financial documents in Confluence, so the OpenAI API is acceptable for triage. The RAG assistant indexes Confluence and Notion, which contain internal financial records and payment data, so the on-prem model is required. The model-agnostic architecture means the application layer is decoupled from the model provider, so the team can swap models without re-implementing the business logic. The two-week pilot should deliver a measured baseline for both workflows: cycle time and error rate for ticket triage, and retrieval accuracy and response quality for the RAG assistant. The pilot should also include a data flow diagram that maps exactly which fields go to the OpenAI API and which stay on the client’s hardware, satisfying the ISO 27001 documentation requirement.

  • AI Automation Glossary for Healthcare and Medtech HR Teams

    Retrieval-Augmented Generation (RAG)

    Retrieval-Augmented Generation (RAG) is a technique that enhances large language models by grounding their responses in a specific, external knowledge base. Instead of relying solely on the model’s pre-trained weights, RAG retrieves relevant documents from a vector database and includes them in the prompt context. This approach is critical for internal knowledge search in healthcare, where accuracy and compliance are paramount. By using RAG, a company can ensure that answers to questions about patient privacy policies or clinical trial protocols are based on the latest internal documentation, reducing the risk of hallucinations and ensuring that the AI provides up-to-date, contextually relevant information. This method allows the AI to act as a knowledgeable assistant that is strictly bound by the company’s own data, making it a reliable tool for both HR and clinical teams.

    pgvector Embeddings Search

    pgvector is an extension for PostgreSQL that enables vector similarity search. It allows developers to store and query high-dimensional vector embeddings directly within a relational database. In the context of internal knowledge search, pgvector is used to index documents from Google Workspace and other sources, converting them into embeddings that can be searched for semantic similarity. This is particularly useful for healthcare and medtech companies that need to maintain strict data governance and ISO 27001 compliance, as it allows the vector database to reside within the same secure, audited environment as other critical data. By using pgvector, organizations can avoid the complexity of managing separate vector databases while still achieving fast and accurate semantic search capabilities, making it a practical choice for scaling AI maturity across departments.

    Workflow Orchestration

    Workflow orchestration is the automated coordination of multiple tasks, systems, and human approvals to achieve a specific business outcome. In AI automation, it involves chaining together document ingestion, vector indexing, LLM inference, and human review steps. For an 11-50 employee healthcare firm, workflow orchestration is essential for managing the complexity of integrating AI into existing processes without disrupting operations. It ensures that data flows correctly between systems, such as from Google Workspace to the RAG pipeline, and that human-in-the-loop approvals are triggered at the right moments. This orchestration layer is what allows the AI system to scale across departments, as it provides a consistent framework for managing different types of workflows, from HR recruiting to clinical documentation, while maintaining compliance and accuracy.

    ISO 27001 Compliance

    ISO 27001 is an international standard for information security management systems (ISMS). It provides a framework for managing sensitive company information so that it remains secure. For healthcare and medtech companies, ISO 27001 compliance is often a requirement for working with partners and patients. When implementing AI automation, the system must be designed to meet these standards, which include strict controls over data access, encryption, and audit logging. This means that the AI system must ensure that patient data and proprietary HR records are processed within these controls, often requiring on-premise or private cloud deployment to prevent data leakage to third-party APIs. Compliance with ISO 27001 is not just a technical requirement but a business enabler, allowing the company to demonstrate its commitment to data security and privacy to stakeholders.

    Human-in-the-Loop (HITL)

    Human-in-the-loop (HITL) is a design pattern where a human is involved in the decision-making process of an AI system. In the context of AI workflow automation, HITL is used to ensure that the AI’s outputs are reviewed and approved by a human before they are finalized or acted upon. This is particularly important in healthcare and HR, where errors can have significant consequences. For example, an AI might draft a response to a policy question or classify a document, but a human must verify the content before it is sent to a candidate or stored in a patient record. HITL helps to maintain trust in the AI system by providing a safety net against errors and ensuring that the AI’s outputs are aligned with the company’s values and compliance requirements. It is a key component of scaling AI maturity across departments, as it allows the company to gradually increase the level of automation while maintaining control and accountability.

    Scaling AI Maturity Across Departments

    AI maturity refers to the level of sophistication and integration of AI capabilities within an organization. Scaling AI maturity across departments involves moving from isolated AI projects to a cohesive, organization-wide AI strategy. For an 11-50 employee healthcare firm, this means expanding the use of AI from a single department, such as HR, to multiple business units, including clinical operations and compliance. This scaling requires a robust infrastructure that can support different types of AI applications, from RAG-based knowledge search to workflow orchestration. It also involves developing the necessary skills and governance frameworks to manage AI across the organization. By scaling AI maturity, the company can achieve greater efficiency, reduce costs, and improve the quality of its services, while also ensuring that its AI initiatives are aligned with its strategic goals and compliance requirements.

    Dedicated AI Team

    A dedicated AI team is a group of specialists who focus exclusively on the development, deployment, and maintenance of AI systems within an organization. Unlike a generalist IT team, a dedicated AI team has the expertise to manage the full lifecycle of AI projects, from initial process audits to ongoing model monitoring and optimization. For a healthcare and medtech company, a dedicated AI team is essential for ensuring that AI initiatives are aligned with the company’s specific needs and compliance requirements. This team is responsible for selecting the right tools and technologies, such as pgvector and RAG, and for integrating them with existing systems like Google Workspace. By having a dedicated AI team, the company can ensure that its AI initiatives are executed efficiently and effectively, while also maintaining the necessary governance and security controls.

  • pgvector RAG and Predictive Scoring for a 12-Person German Fintech

    The Problem: Senior Staff Buried in Routine Queries

    A 12-person fintech in Germany runs on senior engineers and compliance officers who spend 30-40% of their week answering the same questions: “What is our KYC threshold for a new merchant?” “How do we process a chargeback for a card issued in 2019?” “Where is the latest version of our AML policy?” The answers live in Notion, Confluence, and a helpdesk that no one has reorganized since the last product launch. Every query pulls a senior person off their actual work. The cost is not just time—it is the compounding drag on a team that cannot hire a dedicated support layer because the headcount budget is already committed to product and compliance.

    The fix is not a chatbot bolted onto a Slack channel. It is a retrieval-augmented generation (RAG) pipeline that ingests the existing documentation, a predictive scoring model that routes incoming tickets by risk, and a human-in-the-loop approval layer that keeps money-touching actions under human control. The architecture is model-agnostic: OpenAI or Anthropic APIs where quality matters, open-weight models on the client’s own hardware where regulated data cannot leave the building. The integration point is the helpdesk and the documentation platform—Notion or Confluence—via their existing APIs. No new SaaS stack. No rip-and-replace.

    Mechanism: RAG Pipeline and Predictive Scoring

    The pipeline has three stages: ingestion, retrieval, and generation.

    Ingestion. The system pulls documents from Notion or Confluence via their REST APIs. Each document is chunked into 256-512 token segments using a sliding window with 50-token overlap. A sentence-transformer model—BGE-M3 or OpenAI’s text-embedding-3-small—converts each chunk into a 1024-dimensional vector. These vectors store in pgvector, a PostgreSQL extension that adds cosine-similarity search to a standard Postgres instance. For a 10,000-document corpus, the initial index build takes under 5 minutes on a single VPS with 16 GB RAM.

    Retrieval. When a user types a query, the same embedding model converts it to a vector. pgvector returns the top-k (typically k=5) most similar chunks using cosine distance. The query is augmented with metadata filters—document type, last-updated date, access level—so the retrieval respects the team’s existing permission model.

    Generation. The retrieved chunks, the original query, and a system prompt feed into an LLM. The model generates an answer grounded in the retrieved text, with inline citations pointing to the source document and section. For a fintech, the system prompt explicitly instructs the model to flag any answer that touches payment thresholds, AML rules, or contract terms for human review before it reaches the user.

    The predictive scoring model runs in parallel. It is a lightweight classifier—logistic regression or a small feedforward network—trained on historical helpdesk tickets. Features include sender email domain, ticket subject keywords, document type referenced, and time-of-day. The output is a probability score: P(fraud-related), P(AML-related), P(routine). Tickets scoring above 0.7 on fraud or AML route directly to a senior compliance officer. Lower-scoring tickets get an AI-drafted first response for human approval in the helpdesk queue.

    Trade-offs: Model Choice, Chunking, and Approval Scope

    The architect faces three major trade-offs, each with a concrete cost.

    Model choice: cloud API vs. on-premises. OpenAI’s gpt-4o or Anthropic’s claude-3-5-sonnet deliver higher answer quality than open-weight models like Llama 3 70B or Mistral 8x7B. But for a German fintech handling payment data, sending customer names and transaction details to a US-based API may violate internal data-residency policies. The cost of going on-premises: you need a GPU with at least 24 GB VRAM (an A100 or a used RTX 4090 cluster), and the model’s answer quality drops by 10-15% on complex multi-step queries. The mitigation is hybrid: use cloud APIs for internal documentation queries where no customer data is involved, and open-weight models for anything that touches customer PII or payment records.

    Chunking strategy: fixed-size vs. semantic. Fixed 512-token chunks are simple and fast. Semantic chunking—splitting on paragraph boundaries, headings, or natural language breaks—improves retrieval precision by 8-12% but adds complexity to the ingestion pipeline. For a 12-person team, fixed-size chunking with 50-token overlap is the pragmatic default. Semantic chunking becomes worth the engineering time once the corpus exceeds 50,000 documents.

    Human-in-the-loop scope: all responses vs. risk-based. Requiring human approval for every AI-generated response defeats the purpose of automation. The risk-based approach—approve only responses touching money, health data, or contracts—reduces the approval queue by 60-70% while keeping regulatory accountability. The cost: you must define the risk categories precisely and build the routing logic into the helpdesk workflow. For a fintech, the categories are clear: payment processing, AML/KYC, contract terms, and anything involving a customer’s financial data.

    Recommendation: 8-Week Pilot Scope for a 12-Person Fintech

    For a 12-person fintech in Germany, the 8-week pilot follows a fixed scope: one process, one data source, one measurable outcome.

    Weeks 1-2: Process audit. Map the current workflow. Measure baseline cycle time for internal knowledge queries (target: 15-20 minutes per query) and ticket triage error rate (target: 10-15% misclassification). Identify the single highest-ROI process—usually internal knowledge search or ticket triage. Confirm the data source: Notion, Confluence, or both. Document the permission model so the RAG pipeline respects access levels.

    Weeks 3-5: Build. Ingest the documentation corpus into pgvector. Train the predictive scoring model on 6-12 months of historical helpdesk tickets. Build the RAG pipeline with the chosen LLM backend. Integrate with the helpdesk via its API so AI-drafted responses appear in the agent’s queue with confidence scores and source citations.

    Weeks 6-7: Integration and UAT. Connect the pipeline to Notion/Confluence for real-time document updates. Run user acceptance testing with 3-5 senior staff. Measure cycle time and error rate against the baseline. Adjust the risk-based approval thresholds based on UAT feedback.

    Week 8: Go-live and baseline report. Ship the pilot. Produce a before/after report showing cycle time reduction (target: 15-20 min → under 2 min) and error rate change (target: 30-50% reduction in misclassification). The report becomes the business case for rollout to additional processes in subsequent 4-6 week sprints.

    The architecture is deliberately model-agnostic. If the team later migrates from OpenAI to Anthropic, or from cloud to on-premises, the RAG pipeline, embedding model, and scoring logic remain unchanged. The integration point is the LLM API call, not the entire stack.

  • LLM Contract Review for Logistics: pgvector, ISO 27001, and an 8-Week Pilot

    The Problem: Manual Contract Review in a 2,000+ Employee Logistics Firm

    A 2,000+ employee logistics company in the USA processes hundreds of freight forwarding, warehouse, and vendor contracts monthly. Senior staff spend 3-5 hours per contract on manual clause review, with a 15-25% error rate on obligation identification. The cost per contract runs $250-400 in labor, and the cycle time delays onboarding by 5-10 business days. The problem is not a lack of tools but a lack of a structured pipeline that grounds LLM output in the company’s own policy documents and historical precedent while maintaining ISO 27001 audit trails. The pilot must reduce cycle time to under 90 minutes, cut error rates below 5%, and free senior staff for negotiation and exception work within 8 weeks.

    Prerequisites Before Step 1

    Before starting the pilot, confirm the following are in place:

    • API access to the contract repository (e.g., DocuSign, iManage, or a shared drive) and the CRM (Salesforce, HubSpot) where contract metadata lives.
    • Notion or Confluence workspace containing standard clause templates, internal policies, and approval workflows, with read API access enabled.
    • PostgreSQL 15+ with the pgvector extension installed, provisioned on the client’s own infrastructure or a private cloud VPC to satisfy ISO 27001 data residency requirements.
    • LLM API keys for OpenAI (GPT-4o) or Anthropic (Claude 3.5 Sonnet) for the classification and drafting layer, with rate limits and cost caps configured.
    • A named senior reviewer per contract type who will serve as the human-in-the-loop approver during the pilot.
    • Baseline metrics documented: average cycle time, error rate, and cost per contract for the selected contract type over the last 90 days.

    Step 1-3: Build the pgvector Retrieval Layer

    1. Export and chunk policy documents. Pull all standard clause templates and policy statements from Notion or Confluence via their REST APIs. Chunk each document into 200-400 token segments with 50-token overlap. Store the raw text and chunk metadata (source URL, version, last-modified timestamp) in a policy_chunks table in PostgreSQL.

    2. Generate and store embeddings. Use the text-embedding-3-small model (OpenAI) or nomic-embed-text (open-weight, if data cannot leave the building) to generate 1536-dimensional vectors for each chunk. Insert them into a pgvector table with an HNSW index: CREATE INDEX ON policy_chunks USING hnsw (embedding vector_cosine_ops);. Verify index build time is under 5 minutes for 10k chunks.

    3. Build the retrieval function. Write a Python function that takes a contract clause string, embeds it, and queries pgvector for the top-5 most similar policy chunks. Return the chunks with their cosine similarity scores. Set a minimum threshold of 0.75; below this, flag the clause for mandatory human review.

    Step 4-6: LLM Classification and Human Approval

    1. Integrate the LLM classification layer. For each extracted clause, construct a prompt that includes: (a) the clause text, (b) the top-5 retrieved policy chunks with their similarity scores, (c) the contract type and counterparty name. Instruct the model to classify the clause as standard, modified, or non-standard, and to extract all obligations with their source text spans. Use GPT-4o or Claude 3.5 Sonnet with temperature=0.1 for deterministic output.

    2. Add the human approval gate. Route every modified or non-standard clause to the named senior reviewer via a simple web form or Slack integration. The reviewer sees the clause, the retrieved policy context, and the model’s classification. They approve, reject, or edit the classification. Log every decision with a timestamp and reviewer ID for ISO 27001 audit trails.

    3. Implement the secondary verification check. After the LLM extracts obligations, run a second LLM call that verifies each extracted obligation has a direct textual match in the source PDF. If the match score drops below 0.85, log a discrepancy and escalate to a senior reviewer. This catches hallucinated clauses before they reach the approval stage.

    Step 7-9: Orchestration, UAT, and Handoff

    1. Orchestrate the workflow with state tracking. Use Temporal, n8n, or a custom Python state machine to track each contract through stages: ingested, clauses_extracted, classified, pending_approval, approved, signed. Each stage has a timeout (30 minutes for extraction, 4 hours for approval) and a fallback action (escalate to a senior reviewer if approval is not received). Log every state transition with a timestamp, actor, and input/output hashes. Store logs in an append-only table to satisfy ISO 27001 audit requirements.

    2. Run UAT with 20-30 real contracts. Select a mix of standard and complex contracts from the last 90 days. Measure cycle time, error rate, and cost per contract. Compare against the baseline. Target: cycle time under 90 minutes, error rate under 5%, cost per contract under $30. Document all discrepancies and feed them back into the prompt and retrieval thresholds.

    3. Collect ISO 27001 evidence and hand off. Export the audit logs, access control records, and data retention policies. Document the system architecture, API call logs, and encryption configurations. Hand off to the operations team with a runbook covering model version updates, pgvector index maintenance, and escalation paths. The next logical step is to expand the pilot to a second contract type and integrate with the ERP for automated PO generation.

    Common Pitfalls and How to Detect Them

    • Hallucinated clauses. The model invents obligations not present in the source document. Detect via the secondary verification check (match score below 0.85) and the retrieval confidence threshold (below 0.75). Without these guardrails, a single hallucinated indemnity clause can create a $2M+ liability exposure.

    • Stale policy context. The pgvector index contains outdated clause templates because the Notion/Confluence sync failed. Detect by checking the last_synced timestamp in the policy_chunks table and alerting if it exceeds 24 hours. Run a nightly sync job and log failures.

    • Approval bottleneck. Senior reviewers do not respond within the 4-hour window, stalling the pipeline. Detect by monitoring the pending_approval state duration. Escalate to a backup reviewer after 2 hours and log the escalation for process improvement.

    • API cost overrun. Unbounded LLM calls on large contracts (50+ pages) drive API costs above budget. Detect by logging token counts per call and setting a hard cap of 50k tokens per contract. Chunk large contracts and process them in batches.

    • ISO 27001 audit gap. Missing logs for API calls or access control changes. Detect by running a weekly audit log integrity check that verifies every state transition has a corresponding log entry with a hash. Alert on any gaps.

  • Forfis AI Automation Audit and Pilot for Swiss Healthcare and Medtech Operations

    The Problem: Manual Back-Office Work in Swiss Healthcare and Medtech

    You run a 300-person healthcare or medtech company in Switzerland. Your operations team processes 400-600 invoices per month, each taking 12-18 minutes to key into the ERP. Your customer support team handles 150-250 tickets per week, with a median first-response time of 4.2 hours. You want to cut first-response time to under 30 minutes and reduce invoice processing cycle time by 60%, but you cannot send patient-adjacent data to a public cloud API. You need an AI-native operations layer that runs on your own hardware, integrates with your existing ERP and Google Workspace, and ships in 4 weeks. This is the exact scenario Forfis is built for: a fixed-scope pilot on one workflow, measured against a before/after baseline, with human-in-the-loop approval for anything touching money or health data.

    Prerequisites: What You Need Before the Audit Starts

    Before the audit begins, you need four things in place. First, access to your ERP system with read permissions on the invoice module and write permissions on the posting queue. Second, a sample of 50-100 recent invoices in PDF or image format, including at least 10 with line-item errors or missing fields. Third, access to your helpdesk or ticketing system with read permissions on the last 90 days of tickets, including timestamps for first response and resolution. Fourth, a named business owner who can approve scope changes and sign off on the pilot success criteria. You do not need to clean your data before the audit; the audit itself identifies data readiness gaps. You do need to confirm that your IT team can provision a virtual machine or container on your on-premise network for the open-weight model deployment.

    Step 1: Run the Process Audit and Select the Pilot Workflow

    Days 1-5. Forfis reviews your invoice processing workflow end-to-end: how invoices arrive (email, portal, paper), how they are keyed, how errors are handled, and where they sit in the ERP. The deliverable is a process map with cycle time and error rate baselines. You select one workflow for the pilot based on the audit’s prioritization matrix. The pilot scope is fixed: one workflow, one model configuration, one integration point. If you want to automate both invoice processing and customer triage, you run two separate pilots, not one combined engagement.

    Step 2: Deploy the Open-Weight Model on Your On-Premise Hardware

    Days 6-10. Forfis provisions an open-weight model, typically Llama 3 70B or Mistral 8x7B, on your on-premise hardware. The model is fine-tuned on your invoice samples or ticket history, depending on the pilot scope. For invoice processing, the model is trained to extract vendor name, invoice number, line items, tax amounts, and due date from PDF or image input. For customer triage, the model is trained to classify ticket urgency and draft a first response. The fine-tuning dataset is built from your historical data, not synthetic data. You review the model’s output on a holdout set of 20-30 items before it goes live.

    Step 3: Integrate the Agent with Your ERP and Google Workspace

    Days 11-15. Forfis connects the AI agent to your ERP and Google Workspace through their native APIs. For invoice processing, the agent reads the invoice PDF from your email or shared drive, extracts the fields, and posts a draft entry to the ERP posting queue. A human approver reviews the draft in the ERP and clicks approve or reject. For customer triage, the agent reads new tickets from your helpdesk, classifies them, and drafts a first response in Google Workspace. The human agent reviews the draft and sends it. The integration is read-write, so the agent logs its actions in your existing tools without requiring your team to switch platforms.

    Step 4: Run the Pilot in Parallel with Your Existing Process

    Days 16-20. The pilot runs in parallel with your existing process. For invoice processing, the agent processes a subset of invoices, say 20% of the daily volume, while your team continues to process the rest manually. For customer triage, the agent drafts first responses for a subset of tickets, say 30% of the weekly volume, while your team handles the rest. You measure cycle time and error rate for both the agent and the manual process. The success criteria are defined in the audit: for example, a 60% reduction in invoice processing cycle time and a 95% accuracy rate on field extraction. If the agent misses the criteria, Forfis adjusts the model configuration or the integration logic and re-tests.

    Step 5: Validate the Pilot and Roll Out to Full Volume

    Days 21-25. You review the pilot results against the success criteria. If the agent meets the criteria, you proceed to rollout. The rollout expands the agent’s scope from the pilot subset to 100% of the workflow volume. For invoice processing, this means the agent processes all incoming invoices, with human approval still required for anything touching money. For customer triage, this means the agent drafts first responses for all new tickets, with human review before sending. The rollout takes 3-5 business days, during which Forfis monitors the agent’s performance and adjusts thresholds as needed. You do not change your team’s daily workflow; the agent works in the background, and your team approves or rejects its output in the tools they already use.

  • AI Candidate Screening Agent for B2B SaaS Teams in Austria

    The Screening Bottleneck in Small B2B SaaS Teams

    For an 11-50 person B2B SaaS company in Austria, the bottleneck is not a lack of candidates but the time senior staff spend on routine screening. A typical hiring cycle involves parsing 50-100 applications per week, extracting structured data, and drafting first-response emails. This manual work consumes 10-15 hours per week per recruiter, diverting attention from stakeholder alignment and final interviews. The goal is not to replace recruiters but to free them from back-office tasks, enabling them to focus on high-value activities. A conversational agent can handle initial triage, data extraction, and first-response emails, reducing cycle time by 40-60% and error rate by 30-50%. The key is to start with a fixed-scope pilot that measures baseline performance before and after automation, ensuring the investment delivers measurable ROI.

    Architecture: LangGraph Stateful Workflows and RAG

    The agent is built on LangChain and LangGraph, with LangGraph modeling the screening workflow as a stateful graph. This allows for explicit control flow, including human-in-the-loop checkpoints before any action that affects a candidate’s status. The agent uses a retrieval-augmented generation (RAG) approach to access the company’s job descriptions, competency frameworks, and past hiring data. It compares candidate profiles against these criteria, scores them, and flags mismatches. The scoring logic is transparent and auditable, ensuring decisions are based on documented criteria rather than opaque model outputs. For regulated data, the architecture supports open-weight models on the client’s own hardware, ensuring data does not leave the building. This model-agnostic approach allows the company to use OpenAI or Anthropic APIs where quality matters, while maintaining compliance with EU data protection laws.

    Integration with Google Workspace and Existing ATS

    The agent integrates with Google Workspace to read and write emails, access the calendar for scheduling, and retrieve documents from Drive. For candidate screening, the agent parses application emails, extracts structured data (name, experience, skills), and drafts responses. This reduces manual data entry and ensures all candidate interactions are logged in a central system. The integration uses Google’s APIs, avoiding the need to replace existing tools. The agent also connects to the company’s ATS (e.g., Greenhouse, Lever) to update candidate records and trigger next steps. This plug-and-play approach ensures the agent fits into the existing workflow rather than forcing a system change. The result is a seamless reduction in back-office work, with all candidate interactions tracked and auditable.

    Compliance: EU AI Act and GDPR in Austria

    Under the EU AI Act, candidate screening systems are classified as high-risk AI. This requires risk management, data governance, human oversight, and transparency. The agent must operate within a defined scope, and data processing must be documented. Human-in-the-loop design is mandatory for decisions affecting employment, and automated rejections require explicit human review. The system logs all agent actions and human decisions for auditability. In Austria, GDPR also applies, requiring explicit consent and purpose limitation for candidate data. The agent’s scoring logic must be transparent, and candidates must be informed about the use of AI in the screening process. This compliance-first approach ensures the agent meets legal requirements while delivering operational efficiency.

    Two-Week Pilot: Scope, Baseline, and Rollout

    The pilot is scoped to a two-week timeline, assuming the audit is complete and data access is granted. Week 1 focuses on baseline measurement and agent development: the team measures current cycle time and error rate, builds the LangGraph workflow, and sets up the RAG pipeline. Week 2 focuses on integration and human-in-the-loop setup: the agent connects to Google Workspace and the ATS, and the team configures approval steps for high-stakes actions. The pilot ends with a before/after comparison of cycle time and error rate, providing a clear ROI metric. This fixed-scope approach ensures the pilot is deliverable in two weeks and provides a measurable foundation for rollout. The result is a working agent that reduces manual back-office work and frees senior staff for high-value activities.

  • UAE E-commerce: LangGraph Document Extraction and Knowledge Search in Six Months

    The Problem: Routine Work That Should Not Require a Senior Headcount

    A 501-to-2,000-person e-commerce company in the UAE typically runs on a patchwork of Confluence pages, Notion databases, and a CRM that nobody has migrated in three years. The legal and compliance team spends roughly 30 percent of its week pulling product certificates, supplier contracts, and customs declarations out of PDFs, re-keying the data into spreadsheets, and answering the same “where is the compliance file for SKU 4471” question from the operations team. The problem is not a lack of tools; it is that the tools do not talk to each other, and the people who know where things live are the same people who are supposed to be reviewing contracts.

    The fix is not a new platform. It is a fixed-scope integration sprint that inserts an AI layer into the systems you already run. The sprint has a locked scope: one document type, one knowledge-search channel, one measured baseline. It does not replace your CRM, your ERP, or your helpdesk. It plugs into their APIs and adds a retrieval-augmented assistant on top. The architecture is model-agnostic: OpenAI or Anthropic APIs where speed matters, open-weight models on your own hardware where regulated data cannot leave the building. That last point is not optional in the UAE, where data-residency expectations under ISO 27001 Annex A.8.15 and the UAE Data Protection Law mean that a vendor-hosted model is a compliance risk, not just a cost line.

    The Audit: Picking the Workflow That Actually Moves the Needle

    The first two weeks of the engagement are the process audit. The team maps every document that enters the system: supplier invoices, customs declarations, product compliance certificates, internal policy PDFs, and the Confluence pages that hold the answers to “who approved this SKU for the Dubai market?” For each document type, the audit logs the current cycle time, the error rate, and the person who handles it. This is the before/after baseline that the pilot will be measured against.

    The audit also identifies which workflows are worth automating. Not everything is. A document type that appears four times a month and takes eleven minutes to process is not a pilot candidate. The target is a workflow that appears at least 200 times a month, has a measurable error rate above 2 percent, and touches a team that is already at capacity. In a typical UAE e-commerce operation, that is the supplier invoice and the product compliance certificate. The audit output is a one-page scope document that locks the pilot: one document type, one knowledge-search channel, one integration point.

    The scope is fixed. If the team discovers during the build that a second document type would be useful, that is a change request, not a scope expansion. This discipline is what separates an integration sprint from an open-ended consulting engagement, and it is what makes the six-month timeline credible.

    The Build: LangGraph Pipeline with a Human Approval Gate

    The pipeline is built on LangChain for the prompt and tool layer, and LangGraph for the stateful workflow. LangGraph matters here because the document extraction process is not a single call; it is a loop. The model extracts fields from the PDF, a confidence score is computed, and if the score is below 0.85 the item is routed to a human review queue. The human approves, corrects, or rejects. The corrected output is fed back into the training set. LangGraph models this loop as a graph with explicit nodes and edges, so the approval gate is a first-class part of the architecture, not a callback buried in a Python function.

    The knowledge-search assistant uses the same stack. Confluence and Notion both expose REST APIs that return page content as Markdown. The pipeline ingests that content, chunks it by heading, and indexes it in a vector store with metadata: page owner, last-updated date, access level. The LangGraph retrieval node queries the vector store, ranks the top five chunks, and passes them to the LLM for a grounded answer. The answer includes a citation to the source page and a confidence score. For legal and compliance queries, the output is routed to a human reviewer before it reaches the requester. This is the human-in-the-loop default: the model drafts, a person approves anything that touches a contract, a regulation, or a health-data reference.

    The model choice is deferred until the pipeline is working. Weeks two and three use an OpenAI or Anthropic API for speed. Weeks four and five swap to an open-weight model like Llama 3 70B on the client’s own hardware in a UAE data center. The LangGraph interface abstracts the model call, so the swap is a configuration change, not a rewrite.

    The Pilot: Six Weeks, One Document Type, One Measured Baseline

    The pilot runs for six to eight weeks. Week one is the audit and baseline. Weeks two through four are the build: the LangGraph pipeline, the Confluence and Notion API integration, the vector store, and the human review queue. Weeks five through six are the tuning cycle: the team watches the exception rate, adjusts the confidence threshold, and refines the prompt for the document types that are failing. The final two weeks are the measurement: the team compares the pilot’s cycle time and error rate against the baseline from the audit.

    The measurement is not a vanity metric. It is the document that goes to the CFO and the ISO 27001 auditor. The baseline report shows: before the pilot, the supplier invoice took 14 minutes to process and had a 4.2 percent error rate. After the pilot, it takes 3 minutes and the error rate is 0.8 percent. The knowledge-search assistant answered 78 percent of internal queries without a human, and the remaining 22 percent were routed to the review queue with a citation and a confidence score.

    The rollout decision is made at the end of week eight. If the error rate is below 1 percent and the cycle time is below 5 minutes, the pilot graduates to production. The production deployment adds monitoring: the exception rate becomes a KPI in the ISO 27001 operational monitoring plan, and any spike above 3 percent triggers a review of the model or the document format. The managed operation retainer covers the monitoring, the model updates, and the quarterly re-audit of the document types.

    Rollout and Managed Operation: What Happens After the Pilot

    The six-month timeline is not a single sprint. It is a sequence: the audit and pilot in months one and two, the rollout in month three, and the managed operation in months four through six. The rollout is not a big-bang deployment. It is a phased expansion: the first document type goes to production in week nine, the second in week eleven, and the knowledge-search assistant opens to the full team in week thirteen. Each phase has its own baseline measurement and its own exception-rate threshold.

    The managed operation phase is where the engagement stops being a project and starts being a service. The vendor monitors the exception rate, the model performance, and the integration health. If the Confluence API changes its response format, the vendor patches the ingestion layer within 48 hours. If the document format shifts because a new supplier starts sending a different invoice layout, the vendor re-trunes the extraction prompt and re-runs the baseline. The client’s team does not need to hire a data scientist or an ML engineer to keep the system running. That is the point of scaling operations without new hires: the AI layer absorbs the routine work, and the human team focuses on the exceptions and the decisions that actually require judgment.

    The ISO 27001 audit trail is maintained throughout. Every model call, every human approval, every exception routing is logged with a timestamp, the user ID, and the document reference. The logs are stored in the client’s own infrastructure, not in a vendor’s cloud. This is the difference between a system that passes an audit and a system that is built to be audited.

  • AI Workflow Automation vs. Customer Response for E-Commerce in the UAE

    What Is Being Compared

    The two options under comparison are distinct in function, even though both use the same underlying model layer. AI workflow automation targets internal back-office processes: invoice processing, document extraction, and data entry. The goal is to reduce cycle time and error rate in operations and supply chain. Round-the-clock customer response targets external-facing channels: ticket triage, first-response agents, and voice. The goal is to cut first-response time and maintain service levels across time zones. Both options use the OpenAI API as the model layer, integrate with existing tools via API, and ship with a human-in-the-loop approval step. The difference is the workflow being automated and the metric that defines success.

    Criteria for Comparison

    We judge each option against seven criteria that matter to an 11–50 person e-commerce team in the UAE with no specific compliance constraints:

    • Cycle time reduction (internal workflow) vs. first-response time (customer-facing)
    • Error rate (data entry, invoice matching) vs. escalation rate (ticket misclassification)
    • Integration complexity with existing ERP, helpdesk, and documentation tools
    • Human-in-the-loop overhead (approval steps per transaction)
    • Cost per transaction (API call volume, token usage)
    • Time to value within the 8-week fixed-scope pilot
    • Scalability beyond the pilot scope (additional workflows or channels)

    Comparison Table

    Criterion AI Workflow Automation (Invoice Processing) Round-the-Clock Customer Response
    Primary metric Cycle time (hours per invoice) First-response time (minutes per ticket)
    Error rate target <2% mismatch or misclassification <5% misrouted or escalated tickets
    Integration points ERP, accounting software, Notion/Confluence for audit trail Helpdesk, messaging platform, CRM
    Human-in-the-loop Approval before payment or data entry Approval for high-value or sensitive tickets
    API call volume Moderate (one call per invoice) High (one call per ticket, 24/7)
    Time to value in 8 weeks Measurable by week 6 Measurable by week 4
    Scalability Add more invoice types or suppliers Add more channels or languages

    When Each Option Wins

    AI workflow automation wins when the team’s bottleneck is internal: invoice processing is slow, error-prone, and consumes operator time that could go to supply-chain planning. For a 15-person e-commerce team, reducing invoice cycle time from 4 hours to 30 minutes frees up roughly 3.5 operator-hours per invoice. Over 200 invoices per month, that is 700 hours—enough to hire one additional operations analyst or reduce overtime. The fixed-scope pilot delivers a clear before/after baseline on cycle time and error rate, making the business case straightforward.

    Round-the-clock customer response wins when the team’s bottleneck is external: first-response time is high, tickets are piling up, and the team cannot cover all time zones. For an e-commerce business in the UAE serving customers across the Gulf and beyond, a 24/7 AI first-response agent can cut first-response time from 4 hours to 15 minutes. The pilot measures escalation rate and customer satisfaction, and the human-in-the-loop step ensures that high-value or sensitive tickets are routed to a person.

    Recommendation

    For an 11–50 person e-commerce team in the UAE with no specific compliance constraints, AI workflow automation for invoice processing is the stronger first pilot. The reasons are concrete: the workflow is high-volume and repetitive, the success metric (cycle time) is easy to measure, and the human-in-the-loop approval step (before payment) reduces risk. The 8-week timeline is sufficient to audit the process, integrate with the ERP and Notion or Confluence for the audit trail, and deliver a before/after baseline. The OpenAI API is appropriate for the quality of document extraction and classification required. If the pilot meets the target—say, cycle time reduced by 70% and error rate below 2%—the team can scale to additional workflows or add customer-facing automation in a second pilot.

  • AI Workflow Automation vs. Compliance-Safe Rollout for Ticket Triage in B2B SaaS

    What Is Being Compared

    The two options under comparison are AI workflow automation and a compliance-safe AI rollout, both applied to ticket triage and routing in a B2B SaaS company with 2,000+ employees in Switzerland. AI workflow automation refers to the technical layer: an orchestration engine that classifies incoming support tickets, routes them to the correct queue, and drafts a first response using the OpenAI API. It integrates with the existing helpdesk and pulls context from Notion or Confluence via API. The compliance-safe rollout is the delivery and governance layer: a dedicated AI team runs a fixed-scope pilot over 8 weeks, with human-in-the-loop approval on every ticket that touches a customer, and a measured before/after baseline on cycle time and error rate. The two are not alternatives; they are the technical build and the delivery wrapper. The comparison below judges them against the criteria that matter for a 2,000+ employee organization scaling AI across departments.

    Criteria for Judgment

    The following criteria determine which approach fits the scenario. Each is judged against the specific dimensions: B2B SaaS, Switzerland, 2,000+ employees, 8-week timeline, ticket triage and routing, OpenAI API, Notion or Confluence integration, dedicated AI team delivery, and the goal of reducing error rate in the back office.

    • Cycle time reduction: measured from ticket creation to first routed response.
    • Error rate: percentage of misrouted or misclassified tickets.
    • Integration depth: how the AI connects to the helpdesk, Notion/Confluence, and CRM without replacing them.
    • Human-in-the-loop overhead: time a support agent spends approving AI-drafted actions.
    • Timeline feasibility: whether the 8-week window is realistic for pilot and baseline measurement.
    • Scalability across departments: whether the architecture extends to invoice processing, document extraction, and other workflows.
    • Vendor lock-in: whether the model-agnostic design allows swapping OpenAI for an open-weight model if data residency rules change.
    • Cost per ticket: API token cost plus human review time, compared to the current manual triage cost.

    Comparison Table

    Criterion AI Workflow Automation Compliance-Safe Rollout
    Cycle time reduction 40-60% reduction in triage-to-response time Same reduction, but gated by human approval step (adds 5-10 sec per ticket)
    Error rate 30-50% reduction in misrouting Same reduction, with human catch on low-confidence tickets (<0.85)
    Integration depth API connections to helpdesk, Notion/Confluence, CRM Same integrations, plus audit log and approval workflow
    Human-in-the-loop overhead Minimal if confidence threshold is high 5-10 sec per ticket for agent review; scales with ticket volume
    Timeline feasibility 8 weeks for pilot build and baseline 8 weeks includes audit, pilot, tuning, and handover
    Scalability across departments Model-agnostic; new workflows are new integrations Dedicated team runs process audit per department; 2-3 pilots in parallel
    Vendor lock-in OpenAI API; swappable to open-weight model Same; architecture is model-agnostic by design
    Cost per ticket ~EUR 0.02-0.05 in API tokens per ticket Same API cost plus ~EUR 0.10-0.20 in human review time

    Scenario-by-Scenario Verdict

    For a B2B SaaS company in Switzerland with no specific compliance mandate, the AI workflow automation layer is the primary value driver. The OpenAI API handles English-language ticket classification with high accuracy, and the Notion or Confluence integration provides the RAG context for first-response drafting. The 8-week timeline is feasible because the scope is limited to one workflow: ticket triage and routing. The dedicated AI team builds the orchestration, connects the APIs, and runs the pilot. The compliance-safe rollout adds the governance wrapper: human-in-the-loop approval, baseline measurement, and audit logging. For a company with 2,000+ employees, this wrapper is not optional; it is what makes the pilot acceptable to the support leadership and the finance team. The two layers are inseparable in practice: the automation without the rollout wrapper is a demo, not a production system.

    When the company scales across departments, the compliance-safe rollout becomes the scaling mechanism. The dedicated AI team runs a process audit for each new department—invoice processing, document extraction, data entry—and identifies the highest-ROI workflow. The 8-week timeline applies per workflow, not to the entire company. The model-agnostic architecture means each new workflow can use the same orchestration engine, with the OpenAI API for quality-critical tasks and open-weight models on the client’s hardware if a department handles regulated data. The dedicated AI team model ensures continuity: the same team that built the ticket triage pilot runs the next pilot, reducing onboarding friction and maintaining the baseline measurement methodology.

    Recommendation

    The recommendation is to run both layers as a single engagement, not as separate projects. The AI workflow automation is the technical build: an orchestration engine using the OpenAI API that classifies and routes tickets, pulls context from Notion or Confluence, and drafts first responses. The compliance-safe rollout is the delivery and governance wrapper: a dedicated AI team runs the 8-week pilot with human-in-the-loop approval, measures the before/after baseline on cycle time and error rate, and hands over to managed operation. For a 2,000+ employee B2B SaaS company in Switzerland with no compliance constraints, this combined approach is the only one that fits the 8-week timeline and the goal of reducing error rate in the back office. The automation layer delivers the speed and accuracy; the rollout wrapper delivers the trust and the measurement. Neither works without the other. The dedicated AI team owns the technical execution; the client’s support team owns the business outcomes and the human-in-the-loop approval. This split is the standard delivery model for Forfis engagements and is the one that scales across departments without re-architecting the stack.