{"id":141,"date":"2026-10-06T18:59:45","date_gmt":"2026-10-06T18:59:45","guid":{"rendered":"https:\/\/blog.forfis.com\/blog\/ai-support-agent-pilot-checklist-swiss-healthcare\/"},"modified":"2026-10-06T18:59:45","modified_gmt":"2026-10-06T18:59:45","slug":"ai-support-agent-pilot-checklist-swiss-healthcare","status":"publish","type":"post","link":"https:\/\/blog.forfis.com\/blog\/ai-support-agent-pilot-checklist-swiss-healthcare\/","title":{"rendered":"12-Point Checklist: Running a 4-Week AI Support Agent Pilot in Swiss Healthcare"},"content":{"rendered":"<h2>1. Run the process audit and lock the baseline<\/h2>\n<p>Before writing a single line of prompt engineering, the audit must answer three questions: which workflow has the highest volume-to-complexity ratio, which data sources are API-accessible, and which compliance constraints are non-negotiable. For a Swiss healthcare company with no AI in production, the answer is usually ticket triage or first-response drafting on a customer support channel. The audit documents current cycle time (median minutes from ticket open to first human response) and error rate (misrouted or incomplete replies per 100 tickets). These two numbers become the baseline against which the pilot is measured. Without them, the pilot cannot prove ROI. The audit also maps every system the agent will touch\u2014CRM, helpdesk, Notion or Confluence knowledge base\u2014and confirms API credentials, rate limits, and data residency requirements. In Switzerland, FADP and the EU AI Act both apply; the audit flags which fields are personal data, which are health data, and which require human approval before any automated action. The output is a one-page roadmap: one workflow, one integration set, one success metric, four weeks. This document is the contract for the fixed-scope pilot and the reference for every subsequent decision.<\/p>\n<h2>2. Define the fixed-scope pilot boundary<\/h2>\n<p>The pilot scope must be narrow enough to finish in four weeks and broad enough to prove value. For a healthcare and medtech company, the typical scope is a conversational agent that triages incoming support tickets, drafts a first response using the company\u2019s internal knowledge base, and routes the ticket to the right team. The agent does not close tickets, does not touch patient records, and does not send responses without human approval. The knowledge base lives in Notion or Confluence; the agent indexes those spaces via API and retrieves relevant passages to ground every draft. The CRM and helpdesk integrations are read-write for ticket metadata and read-only for customer history. The Anthropic Claude API handles classification and drafting; the model is selected for its instruction-following quality and context window, not for cost. The architecture is model-agnostic: if the client later moves to an open-weight model on local hardware for data residency reasons, the prompt layer and integration layer remain unchanged. The pilot ships with a dashboard showing cycle time, error rate, and human override rate, updated daily. At week four, the team compares the pilot numbers against the audit baseline and makes a go\/no-go decision on rollout.<\/p>\n<h2>3. Configure EU AI Act and Swiss FADP compliance gates<\/h2>\n<p>The EU AI Act, effective in phases from 2025, requires transparency for AI systems that interact with humans. Article 50 mandates that users be informed they are interacting with an AI, unless it is obvious from context. For a healthcare support agent, this means the first message must state that the response is AI-drafted and subject to human review. The Act also classifies systems that make decisions affecting health as high-risk under Article 6, but a triage-and-draft agent that does not diagnose, prescribe, or alter treatment plans falls outside that category. Still, the agent must not process health data without a legal basis under GDPR and Swiss FADP. The pilot configuration includes a data classification layer: fields tagged as health data are routed to a human approver before any action. The agent\u2019s system prompt explicitly forbids it from making medical claims, interpreting test results, or advising on treatment. Every response is logged with the model version, prompt hash, and retrieval context for auditability. The compliance checklist is signed off by the client\u2019s data protection officer before the pilot goes live, and the log retention period matches the client\u2019s regulatory requirement, typically 12 months for healthcare records in Switzerland.<\/p>\n<h2>4. Build the retrieval layer over Notion or Confluence<\/h2>\n<p>The agent\u2019s value depends on retrieval quality. The knowledge base in Notion or Confluence must be structured so the agent can find the right passage in under 200 ms. Before the pilot, the team runs a retrieval audit: take 50 real support tickets from the past quarter, identify the correct knowledge base article for each, and measure how often a vector search over the raw document text returns that article in the top three results. If the hit rate is below 80%, the knowledge base needs restructuring before the agent is built. Concretely, this means splitting long pages into discrete, self-contained sections, adding metadata tags (product, issue type, severity), and removing deprecated content. The retrieval pipeline uses a hybrid approach: dense vector embeddings for semantic matching and BM25 for exact keyword hits, with a reranking step using the Claude API to score the top ten candidates. The agent\u2019s system prompt instructs it to cite the specific knowledge base section in every draft, so the human approver can verify the source. If the retrieval confidence score falls below a threshold the team sets during the audit, the agent flags the ticket for manual handling rather than drafting a potentially wrong response. This guardrail is non-negotiable in a healthcare context.<\/p>\n<h2>5. Measure cycle time, error rate, and override rate daily<\/h2>\n<p>The pilot runs for four weeks with a daily standup and a weekly metrics review. The team tracks three numbers every day: median cycle time from ticket open to first human-approved response, error rate (tickets requiring rework after approval), and human override rate (percentage of drafts the approver rejects or significantly edits). The audit baseline from step one is the reference. A successful pilot shows at least a 30% reduction in cycle time and a 20% reduction in error rate, with an override rate below 15% by week three. If the override rate stays above 25%, the team investigates: is the retrieval missing the right article, is the prompt too vague, or is the knowledge base outdated? The fix is applied within 48 hours and the metrics are re-measured. The pilot also includes a shadow mode for the first three days: the agent drafts responses but does not send them; the human approver compares the draft against what they would have written. This calibrates the prompt and the retrieval thresholds before the agent goes live. At the end of week four, the team produces a one-page report: baseline vs. pilot numbers, override rate trend, top five failure modes, and a recommendation on rollout scope. The report is the input to the next engagement, not a marketing document.<\/p>\n<h2>6. Maintain the checklist and the agent after go-live<\/h2>\n<p>The pilot is not a one-and-done deliverable. The knowledge base in Notion or Confluence changes weekly; new product releases, policy updates, and support macros all alter the retrieval landscape. The team schedules a monthly retrieval audit: take 20 new tickets, measure the hit rate, and restructure sections if the rate drops below 80%. The prompt layer is versioned in a repository with a changelog; every change is tested against a fixed set of 30 evaluation tickets before deployment. The compliance log is reviewed quarterly by the data protection officer to confirm that no health data was processed without approval and that the AI transparency notice is still present in every first response. The model provider\u2019s terms of service and the EU AI Act\u2019s obligations are re-checked at each quarterly review, because both evolve. The team also maintains a runbook for model degradation: if the Claude API\u2019s response quality drops due to a provider-side change, the runbook specifies the fallback\u2014switch to the open-weight model on local hardware, re-run the evaluation set, and deploy within 24 hours. The checklist itself is stored in the same Notion or Confluence space the agent indexes, so the team can search for it the same way the agent searches for support articles. This keeps the maintenance process visible and auditable.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A 12-item checklist for running a 4-week fixed-scope pilot of an Anthropic Claude-powered support agent in a Swiss healthcare company, covering EU AI Act compliance, Notion integration, and measurable baselines.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"12-Point Checklist: Running a 4-Week AI Support Agent Pilot in Swiss Healthcare","rank_math_description":"A 12-item checklist for running a 4-week fixed-scope pilot of an Anthropic Claude-powered support agent in a Swiss healthcare company, covering EU AI Act compliance, Notion integration, and measurable baselines.","rank_math_focus_keyword":"replace manual data entry internal knowledge search","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","_yoast_wpseo_focuskw":"","pll_lang":"en","geo_jsonld":"{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-support-agent-pilot-checklist-swiss-healthcare\/#article\",\"@type\":\"Article\",\"author\":{\"@id\":\"https:\/\/blog.forfis.com#org\"},\"dateModified\":\"2026-10-05T23:48:06.099239391+00:00\",\"datePublished\":\"2026-10-05T23:48:06.099239391+00:00\",\"description\":\"A 12-item checklist for running a 4-week fixed-scope pilot of an Anthropic Claude-powered support agent in a Swiss healthcare company, covering EU AI Act compliance, Notion integration, and measurable baselines.\",\"headline\":\"12-Point Checklist: Running a 4-Week AI Support Agent Pilot in Swiss Healthcare\",\"inLanguage\":\"en\",\"keywords\":[\"No AI in Production Yet\",\"Anthropic Claude API\",\"Conversational Agent\",\"Customer Support\",\"11-50\",\"EU AI Act\",\"Fixed-Scope Pilot\",\"Healthcare and Medtech\",\"Notion or Confluence\",\"English\",\"Replace Manual Data Entry\",\"Switzerland\",\"4 weeks\",\"Internal Knowledge Search\"],\"mainEntityOfPage\":\"https:\/\/blog.forfis.com\/blog\/ai-support-agent-pilot-checklist-swiss-healthcare\/\",\"publisher\":{\"@id\":\"https:\/\/blog.forfis.com#org\"}},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-support-agent-pilot-checklist-swiss-healthcare\/#faq\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The EU AI Act classifies a support chatbot as a limited-risk system under Article 50, requiring transparency about AI interaction. If the agent touches patient data, GDPR and Swiss FADP apply. Forfis builds a human-in-the-loop gate so no automated action touches money, health data, or contracts without approval.\"},\"name\":\"What compliance obligations apply to a healthcare support agent in Switzerland?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"A 4-week fixed-scope pilot covers one workflow end-to-end: audit, build, integrate, and measure. It ships with a before\/after baseline on cycle time and error rate. Rollout to additional channels or teams follows as a separate engagement once the pilot validates ROI.\"},\"name\":\"How does a 4-week fixed-scope pilot differ from a full rollout?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Forfis uses Anthropic Claude API where quality matters and open-weight models on client hardware where regulated data cannot leave the building. The architecture is model-agnostic, so switching providers later does not require rebuilding the agent or its integrations.\"},\"name\":\"Why choose Anthropic Claude over open-weight models for a healthcare agent?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The agent drafts responses and classifies tickets; a human approves anything touching money, health data, or a contract. Every pilot ships with a measured baseline on cycle time and error rate so the team can quantify improvement against the pre-automation state.\"},\"name\":\"How does human-in-the-loop work in practice for a support agent?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The agent plugs into existing CRMs, ERPs, helpdesks, and messaging through their APIs rather than replacing them. For knowledge search, it indexes Notion or Confluence spaces and retrieves relevant passages to ground responses, reducing manual data entry and copy-paste work.\"},\"name\":\"How does the agent integrate with Notion or Confluence without replacing them?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The process audit identifies workflows worth automating based on volume, error rate, and cycle time. For a company with no AI in production, the audit also maps data readiness, API access, and compliance constraints before selecting the first pilot workflow.\"},\"name\":\"What does the AI process audit cover for a company with no AI in production yet?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"For an 11-50 person team, the pilot targets one high-volume, low-complexity workflow\u2014typically ticket triage or first-response drafting. The fixed scope keeps the team focused: one integration, one measured outcome, and a clear go\/no-go decision at week four.\"},\"name\":\"How do we scope a pilot for a team of 11-50 people?\"}]},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-support-agent-pilot-checklist-swiss-healthcare\/#breadcrumbs\",\"@type\":\"BreadcrumbList\",\"itemListElement\":[{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\",\"name\":\"Home\",\"position\":1},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/\",\"name\":\"Blog\",\"position\":2},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/ai-support-agent-pilot-checklist-swiss-healthcare\/\",\"name\":\"12-Point Checklist: Running a 4-Week AI Support Agent Pilot in Swiss Healthcare\",\"position\":3}]},{\"@id\":\"https:\/\/blog.forfis.com#org\",\"@type\":\"Organization\",\"name\":\"Forfis\",\"url\":\"https:\/\/blog.forfis.com\"}]}","geo_content_hash":"b89ff6bdd5fa6ae66cea6d11b3c42f2660e49a34b0b14c060acd40105d22cb52","footnotes":""},"categories":[45],"tags":[47,73,43],"class_list":["post-141","post","type-post","status-publish","format-standard","hentry","category-healthcare-and-medtech","tag-internal-knowledge-search","tag-replace-manual-data-entry","tag-switzerland"],"_links":{"self":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/141","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/comments?post=141"}],"version-history":[{"count":0,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/141\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/media?parent=141"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/categories?post=141"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/tags?post=141"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}