1. Run the process audit and lock the baseline
Before writing a single line of prompt engineering, the audit must answer three questions: which workflow has the highest volume-to-complexity ratio, which data sources are API-accessible, and which compliance constraints are non-negotiable. For a Swiss healthcare company with no AI in production, the answer is usually ticket triage or first-response drafting on a customer support channel. The audit documents current cycle time (median minutes from ticket open to first human response) and error rate (misrouted or incomplete replies per 100 tickets). These two numbers become the baseline against which the pilot is measured. Without them, the pilot cannot prove ROI. The audit also maps every system the agent will touch—CRM, helpdesk, Notion or Confluence knowledge base—and confirms API credentials, rate limits, and data residency requirements. In Switzerland, FADP and the EU AI Act both apply; the audit flags which fields are personal data, which are health data, and which require human approval before any automated action. The output is a one-page roadmap: one workflow, one integration set, one success metric, four weeks. This document is the contract for the fixed-scope pilot and the reference for every subsequent decision.
2. Define the fixed-scope pilot boundary
The pilot scope must be narrow enough to finish in four weeks and broad enough to prove value. For a healthcare and medtech company, the typical scope is a conversational agent that triages incoming support tickets, drafts a first response using the company’s internal knowledge base, and routes the ticket to the right team. The agent does not close tickets, does not touch patient records, and does not send responses without human approval. The knowledge base lives in Notion or Confluence; the agent indexes those spaces via API and retrieves relevant passages to ground every draft. The CRM and helpdesk integrations are read-write for ticket metadata and read-only for customer history. The Anthropic Claude API handles classification and drafting; the model is selected for its instruction-following quality and context window, not for cost. The architecture is model-agnostic: if the client later moves to an open-weight model on local hardware for data residency reasons, the prompt layer and integration layer remain unchanged. The pilot ships with a dashboard showing cycle time, error rate, and human override rate, updated daily. At week four, the team compares the pilot numbers against the audit baseline and makes a go/no-go decision on rollout.
3. Configure EU AI Act and Swiss FADP compliance gates
The EU AI Act, effective in phases from 2025, requires transparency for AI systems that interact with humans. Article 50 mandates that users be informed they are interacting with an AI, unless it is obvious from context. For a healthcare support agent, this means the first message must state that the response is AI-drafted and subject to human review. The Act also classifies systems that make decisions affecting health as high-risk under Article 6, but a triage-and-draft agent that does not diagnose, prescribe, or alter treatment plans falls outside that category. Still, the agent must not process health data without a legal basis under GDPR and Swiss FADP. The pilot configuration includes a data classification layer: fields tagged as health data are routed to a human approver before any action. The agent’s system prompt explicitly forbids it from making medical claims, interpreting test results, or advising on treatment. Every response is logged with the model version, prompt hash, and retrieval context for auditability. The compliance checklist is signed off by the client’s data protection officer before the pilot goes live, and the log retention period matches the client’s regulatory requirement, typically 12 months for healthcare records in Switzerland.
4. Build the retrieval layer over Notion or Confluence
The agent’s value depends on retrieval quality. The knowledge base in Notion or Confluence must be structured so the agent can find the right passage in under 200 ms. Before the pilot, the team runs a retrieval audit: take 50 real support tickets from the past quarter, identify the correct knowledge base article for each, and measure how often a vector search over the raw document text returns that article in the top three results. If the hit rate is below 80%, the knowledge base needs restructuring before the agent is built. Concretely, this means splitting long pages into discrete, self-contained sections, adding metadata tags (product, issue type, severity), and removing deprecated content. The retrieval pipeline uses a hybrid approach: dense vector embeddings for semantic matching and BM25 for exact keyword hits, with a reranking step using the Claude API to score the top ten candidates. The agent’s system prompt instructs it to cite the specific knowledge base section in every draft, so the human approver can verify the source. If the retrieval confidence score falls below a threshold the team sets during the audit, the agent flags the ticket for manual handling rather than drafting a potentially wrong response. This guardrail is non-negotiable in a healthcare context.
5. Measure cycle time, error rate, and override rate daily
The pilot runs for four weeks with a daily standup and a weekly metrics review. The team tracks three numbers every day: median cycle time from ticket open to first human-approved response, error rate (tickets requiring rework after approval), and human override rate (percentage of drafts the approver rejects or significantly edits). The audit baseline from step one is the reference. A successful pilot shows at least a 30% reduction in cycle time and a 20% reduction in error rate, with an override rate below 15% by week three. If the override rate stays above 25%, the team investigates: is the retrieval missing the right article, is the prompt too vague, or is the knowledge base outdated? The fix is applied within 48 hours and the metrics are re-measured. The pilot also includes a shadow mode for the first three days: the agent drafts responses but does not send them; the human approver compares the draft against what they would have written. This calibrates the prompt and the retrieval thresholds before the agent goes live. At the end of week four, the team produces a one-page report: baseline vs. pilot numbers, override rate trend, top five failure modes, and a recommendation on rollout scope. The report is the input to the next engagement, not a marketing document.
6. Maintain the checklist and the agent after go-live
The pilot is not a one-and-done deliverable. The knowledge base in Notion or Confluence changes weekly; new product releases, policy updates, and support macros all alter the retrieval landscape. The team schedules a monthly retrieval audit: take 20 new tickets, measure the hit rate, and restructure sections if the rate drops below 80%. The prompt layer is versioned in a repository with a changelog; every change is tested against a fixed set of 30 evaluation tickets before deployment. The compliance log is reviewed quarterly by the data protection officer to confirm that no health data was processed without approval and that the AI transparency notice is still present in every first response. The model provider’s terms of service and the EU AI Act’s obligations are re-checked at each quarterly review, because both evolve. The team also maintains a runbook for model degradation: if the Claude API’s response quality drops due to a provider-side change, the runbook specifies the fallback—switch to the open-weight model on local hardware, re-run the evaluation set, and deploy within 24 hours. The checklist itself is stored in the same Notion or Confluence space the agent indexes, so the team can search for it the same way the agent searches for support articles. This keeps the maintenance process visible and auditable.
Leave a Reply