The problem: manual ticket triage at 500 tickets per week
A 120-person US healthcare operations team handles 500+ support tickets per week across billing, clinical queries, and supply chain issues. Every ticket lands in a shared queue, a human reads it, decides the category, and routes it to the right specialist. Cycle time averages 4.2 hours; misrouting rate sits at 12%. The team cannot hire more triage staff without breaking the operating budget, and the current process does not scale with ticket volume. The problem is not a lack of tools — it is that the routing decision is manual, slow, and inconsistent. The fix is an AI agent that classifies and routes tickets automatically, with a human approval gate for anything touching PHI, billing, or contracts. The delivery vehicle is an 8-week fixed-scope pilot built on LangChain and LangGraph, integrated into the team’s existing Slack workspace, and measured against a before/after baseline on cycle time and error rate.
Prerequisites: what you need before week 1
Before the pilot begins, you need four things in place. First, a process audit that documents the current triage workflow: which queues exist, what categories are used, what the routing rules are, and where the bottlenecks sit. Forfis runs this audit in week 1 and produces a one-page map of the workflow. Second, API access to your ticketing system (Zendesk, Freshdesk, or equivalent) and to Slack or Microsoft Teams. You need read/write scopes for ticket creation, status updates, and channel posting. Third, a HIPAA compliance review: confirm whether the ticket data contains PHI, identify which fields are sensitive, and determine whether a BAA is required with any third-party LLM provider. Fourth, a baseline measurement: pull 2 weeks of historical ticket data and record cycle time (creation to first routed response) and misrouting rate. This baseline is the number the pilot must beat.
Step 1: Run the process audit and lock the scope
Week 1 is the process audit. Forfis maps the current triage workflow end-to-end: ticket intake, category assignment, routing rules, escalation paths, and resolution. The output is a one-page workflow diagram and a list of the top 5 routing rules that account for 80% of ticket volume. You review this map and confirm the scope: which ticket categories the pilot will cover, which queues it will route to, and which fields are PHI. This step prevents scope creep later. The audit also identifies the integration points: which API endpoints the agent will call, what authentication method your ticketing system uses, and whether Slack or Teams is the primary notification channel. You sign off on the scope document before week 2 begins.
Step 2: Design the LangGraph agent with human-in-the-loop gates
Weeks 2-3 are the agent design and build. Forfis constructs the triage agent using LangGraph as the state machine and LangChain for LLM abstraction. The graph has four nodes: classify (LLM assigns a category from your taxonomy), route (conditional branch sends the ticket to the correct queue), approve (human-in-the-loop gate for PHI, billing, or contract tickets), and notify (posts the routing decision to Slack or Teams). The classify node uses a structured output schema so the LLM returns a JSON object with category, confidence, and routing_target. The approve node pauses execution and sends an approval request to the designated human via Slack. For regulated data, the LLM runs on your own hardware using an open-weight model (Llama 3 70B or Mistral 7B) to keep PHI inside your network. For non-PHI classification, an OpenAI or Anthropic API call is acceptable. The agent is tested against 200 historical tickets before the pilot goes live.
Step 3: Integrate with Slack or Teams and run the pilot
Weeks 4-5 are the pilot build and integration. The agent connects to your ticketing system via its REST API: it reads new tickets, classifies them, and writes the routing decision back to the ticket’s status field. The Slack or Teams integration posts a message to the operations channel with the ticket ID, assigned category, routing target, and confidence score. For multilingual support, the agent detects the ticket language using a lightweight classifier (fasttext or the LLM itself) and processes the ticket in that language. The routing rules are the same regardless of language; only the classification prompt is localized. The human approval gate is configured so that any ticket with a confidence score below 0.85, or any ticket tagged as PHI, billing, or contract, requires a human to click “Approve” or “Reject” in Slack before the routing is executed. The pilot runs on a subset of tickets — typically 20% of volume — so the team can compare AI-routed tickets against human-routed ones side by side.
Step 4: Measure the pilot against the baseline
Weeks 6-7 are pilot operation and baseline comparison. The agent runs on the 20% pilot subset for 2 weeks. Forfis tracks three metrics daily: cycle time (creation to first routed response), misrouting rate (tickets sent to the wrong queue), and human override rate (percentage of AI decisions that a human rejected or modified). At the end of week 7, Forfis produces a comparison report: baseline vs. pilot on all three metrics. A typical result for a 120-person healthcare operations team is a 45% reduction in cycle time (from 4.2 hours to 2.3 hours) and a 50% reduction in misrouting (from 12% to 6%). The human override rate should be below 15% by the end of the pilot; if it is higher, the classification prompts need tuning before rollout. The report also flags any tickets where the agent failed to detect PHI or misclassified a clinical query as a billing issue — these are the edge cases that need prompt refinement.
Step 5: Go/no-go review and rollout plan
Week 8 is the go/no-go review. You and Forfis sit down with the comparison report and decide: does the pilot meet the success criteria? The criteria are defined in the scope document from week 1 — typically a 40%+ reduction in cycle time and a 50%+ reduction in misrouting, with a human override rate below 15%. If the pilot meets the criteria, the next step is a rollout plan: expand the agent to 100% of ticket volume, add the remaining ticket categories, and set up ongoing monitoring. If the pilot misses the criteria, Forfis identifies the specific failure modes (usually prompt gaps on edge-case categories or integration latency) and proposes a 2-week remediation sprint before re-running the pilot. The rollout plan includes a managed operation phase: Forfis monitors the agent’s performance, tunes prompts as new ticket patterns emerge, and handles model updates. The architecture is model-agnostic, so if a new open-weight model outperforms the current one, the swap is a configuration change, not a rebuild.