{"id":483,"date":"2026-10-06T19:00:43","date_gmt":"2026-10-06T19:00:43","guid":{"rendered":"https:\/\/blog.forfis.com\/blog\/us-insurer-voice-agent-shipment-status-4-week-pilot\/"},"modified":"2026-10-06T19:00:43","modified_gmt":"2026-10-06T19:00:43","slug":"us-insurer-voice-agent-shipment-status-4-week-pilot","status":"publish","type":"post","link":"https:\/\/blog.forfis.com\/blog\/us-insurer-voice-agent-shipment-status-4-week-pilot\/","title":{"rendered":"How a 2,400-Person US Insurer Cut Shipment-Status Call Time by 67% in 4 Weeks"},"content":{"rendered":"<h2>Background: A 2,400-Person US Insurer with a 18,000-Call Monthly Queue<\/h2>\n<p>This case study is a composite drawn from patterns observed across multiple insurance and insurtech engagements. No named customer is represented. The company profile, metrics, and timeline reflect the median outcome from a cohort of similar deployments, not a single client.<\/p>\n<p>The company is a mid-size US property and casualty insurer with 2,400 employees, headquartered in Columbus, Ohio. It writes personal auto, home, and commercial lines. The customer support operation handles roughly 18,000 inbound calls per month, of which 60-70% are status inquiries: \u201cWhere is my claim check?\u201d, \u201cHas my replacement part shipped?\u201d, \u201cWhat is the ETA on my repair?\u201d The existing stack includes a Genesys Cloud contact center, a custom TMS built on PostgreSQL with a REST API, and a Salesforce CRM. The support team is staffed 24\/7 across three shifts, with an average handle time of 4 minutes 12 seconds for status calls and a first-contact resolution rate of 71%.<\/p>\n<h2>Challenge: 60% of Calls Were Status Checks, and the 4-Week Deadline Was Non-Negotiable<\/h2>\n<p>The operational pressure was threefold. First, the support team was at 94% utilization during peak hours (9 AM-1 PM ET), with average wait times exceeding 6 minutes. Second, the company had committed to a GDPR-aligned data handling policy for its US operations after a 2024 regulatory review, which meant any new system touching caller PII had to keep data on-premises or in a US-only cloud region with explicit consent logging. Third, the CFO had set a 4-week deadline for a pilot that would demonstrate measurable cycle-time reduction before the Q3 budget cycle closed. The specific need was to replace the manual data-entry step where agents typed shipment IDs into the TMS, waited for a status, and read it back. That step alone consumed 55-70 seconds of every status call.<\/p>\n<h2>Approach: Self-Hosted Voice Agent on LangGraph with a Fixed 4-Week Pilot Scope<\/h2>\n<p>The dedicated AI team consisted of one ML engineer, one full-stack developer, one product manager, and one QA specialist, embedded with the client\u2019s IT and support operations teams. The architecture was model-agnostic by design: the LLM layer ran on a self-hosted Llama-3-70B instance on the client\u2019s on-premises GPU cluster, the ASR used Whisper-large-v3 fine-tuned on insurance terminology, and the TTS used a fine-tuned Coqui TTS model. Orchestration was built on <strong>LangGraph<\/strong>, which managed the conversation state machine: greeting, identity verification, intent classification, TMS query, status readout, and transfer-to-human. The TMS integration used the existing REST API with webhook callbacks for status changes. No proprietary SaaS voice platform was used. The pilot scope was fixed: one carrier, one status type (shipment ETA), one language (English), and a hard boundary that the agent would not accept payment, modify policy terms, or initiate claims.<\/p>\n<h2>Outcome: 67% Cycle-Time Reduction and 88% First-Contact Resolution in 4 Weeks<\/h2>\n<p>The pilot ran for 4 weeks, with the agent handling 15% of inbound status calls in week 2, 30% in week 3, and 50% in week 4. Baseline metrics were captured in week 1 from 200 sampled calls in the human queue. By the end of week 4, the agent\u2019s average handle time for status queries was 82 seconds, compared to the human baseline of 252 seconds \u2014 a 67% reduction. First-contact resolution for status-only calls reached 88%, up from the 71% human baseline. The error rate on status readout was 1.4%, below the 2% threshold. The agent transferred 22% of calls to humans, primarily for claim disputes and policy changes. The client\u2019s support team reported that the 15-30% of calls absorbed by the agent freed agents to handle complex cases, reducing average wait time during peak hours from 6 minutes to under 3 minutes. The pilot met all three KPI targets for 5 consecutive business days before the client approved rollout to 100% of status calls.<\/p>\n<h2>Lessons for Similar Teams Scaling Voice Automation Across Departments<\/h2>\n<ul>\n<li><strong>Fix the TMS API before building the agent.<\/strong> The client\u2019s TMS REST API had undocumented rate limits (50 requests\/minute) and inconsistent status codes across three carrier integrations. Two days of the 4-week timeline were consumed normalizing the API response schema. If the API is not stable, the agent will inherit the inconsistency and the error rate will exceed the threshold.<\/li>\n<li><strong>Identity verification is the single biggest failure point.<\/strong> The agent\u2019s confidence in caller identity dropped below 90% when callers provided partial policy numbers or used different names than on file. The LangGraph state machine needed a fallback path that gracefully degraded to a human transfer rather than guessing. Budget time for this edge case.<\/li>\n<li><strong>GDPR compliance is an architecture decision, not a checkbox.<\/strong> Keeping ASR and LLM inference on-premises was non-negotiable. The client\u2019s legal team required that no raw audio or PII left the building. This constraint shaped the entire stack selection and added 3 days of infrastructure setup.<\/li>\n<li><strong>The 4-week timeline is only realistic with a fixed scope.<\/strong> Expanding the pilot to multi-carrier, multi-language, or claim-initiation use cases would have pushed the timeline to 7-9 weeks. The client\u2019s commitment to a single use case was the critical enabler.<\/li>\n<li><strong>Human-in-the-loop is not optional for regulated industries.<\/strong> The agent\u2019s hard boundary on payment, policy modification, and claim initiation was enforced in the LangGraph state machine, not in the prompt. Model-level instructions are not a compliance control.<\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>A 2,400-person US insurer cut shipment-status call handling time by 62% in 4 weeks using a self-hosted voice agent built on LangGraph and Llama-3. A composite case study with concrete metrics and GDPR controls.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"How a 2,400-Person US Insurer Cut Shipment-Status Call Time by 67% in 4 Weeks","rank_math_description":"A 2,400-person US insurer cut shipment-status call handling time by 62% in 4 weeks using a self-hosted voice agent built on LangGraph and Llama-3. A composite case study with concrete metrics and GDPR controls.","rank_math_focus_keyword":"replace manual data entry order and shipment status updates","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","_yoast_wpseo_focuskw":"","pll_lang":"en","geo_jsonld":"{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@id\":\"https:\/\/blog.forfis.com\/blog\/us-insurer-voice-agent-shipment-status-4-week-pilot\/#article\",\"@type\":\"Article\",\"author\":{\"@id\":\"https:\/\/blog.forfis.com#org\"},\"dateModified\":\"2026-10-06T00:01:22.957187095+00:00\",\"datePublished\":\"2026-10-06T00:01:22.957187095+00:00\",\"description\":\"A 2,400-person US insurer cut shipment-status call handling time by 62% in 4 weeks using a self-hosted voice agent built on LangGraph and Llama-3. A composite case study with concrete metrics and GDPR controls.\",\"headline\":\"How a 2,400-Person US Insurer Cut Shipment-Status Call Time by 67% in 4 Weeks\",\"inLanguage\":\"en\",\"keywords\":[\"Scaling Across Departments\",\"LangChain and LangGraph\",\"Voice Agent\",\"Customer Support\",\"2000+\",\"GDPR\",\"Dedicated AI Team\",\"Insurance and Insurtech\",\"Custom REST API and Webhooks\",\"English\",\"Replace Manual Data Entry\",\"USA\",\"4 weeks\",\"Order and Shipment Status Updates\"],\"mainEntityOfPage\":\"https:\/\/blog.forfis.com\/blog\/us-insurer-voice-agent-shipment-status-4-week-pilot\/\",\"publisher\":{\"@id\":\"https:\/\/blog.forfis.com#org\"}},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/us-insurer-voice-agent-shipment-status-4-week-pilot\/#faq\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The voice agent handles the full conversation: it greets the caller, verifies identity using the last four digits of the policy number and date of birth, queries the TMS via REST, and reads back the status. If the caller asks to file a claim, change a beneficiary, or disputes the status, the agent transfers to a human queue. The agent never accepts payment or modifies policy terms. This boundary is hard-coded in the LangGraph state machine, not left to model discretion.\"},\"name\":\"What does the voice agent actually do versus what a human still handles?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The agent runs on a dedicated VPC in us-east-1. Audio is transcribed by an on-premises ASR model (Whisper-large-v3 fine-tuned on insurance terminology) so raw voice data never leaves the client's infrastructure. The LLM inference runs on a self-hosted Llama-3-70B instance. Only the structured status payload (shipment ID, status code, ETA) is sent to the TMS via the existing REST API. No PII is logged in the LLM context window; the conversation log is encrypted at rest with AES-256 and access-controlled per GDPR Article 32.\"},\"name\":\"How is GDPR compliance maintained when a voice agent processes caller data?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The 4-week timeline assumes the TMS REST API is documented and stable, the client has a single primary use case (shipment status), and the dedicated team is fully staffed from day one. If the TMS API requires custom authentication, rate-limiting workarounds, or the use case expands to include claim initiation, add 2-3 weeks. The pilot scope is fixed: one carrier, one status type, one language. Expanding to multi-carrier or multi-language support is a separate phase.\"},\"name\":\"Is the 4-week timeline realistic for a 2,000+ employee insurance company?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The voice agent runs 24\/7 on the client's infrastructure. The dedicated AI team provides 9 AM-6 PM ET coverage for monitoring, model updates, and incident response. The client's internal IT team handles infrastructure uptime. For a 2,000+ employee company, the dedicated team typically includes one ML engineer, one full-stack developer, one product manager, and one QA specialist. The client retains ownership of all code, models, and data; Forfis does not retain any IP.\"},\"name\":\"What does the dedicated AI team look like during and after the 4-week pilot?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The agent uses a self-hosted Llama-3-70B model for the LLM layer, Whisper-large-v3 for ASR, and a fine-tuned TTS model for voice output. The orchestration layer is LangGraph, which manages the state machine for the conversation flow. The TMS integration uses the client's existing REST API with webhook callbacks for status changes. No proprietary SaaS voice platform is used; all components are open-weight or self-hosted to keep data on-premises.\"},\"name\":\"What models and frameworks does the voice agent use?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot measures three KPIs: average handle time (target: under 90 seconds for status queries), first-contact resolution rate (target: 85%+ for status-only calls), and error rate on status readout (target: under 2%). The baseline is captured during week 1 by sampling 200 calls from the existing human queue. The agent is considered ready for rollout when it meets all three targets for 5 consecutive business days. If error rate exceeds 2%, the agent is pulled from the queue and the LangGraph state machine is debugged before re-deployment.\"},\"name\":\"How do we measure success in the 4-week pilot?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The agent does not replace the human queue; it absorbs the status-check volume that currently occupies 60-70% of inbound calls. Human agents handle claims, policy changes, disputes, and complex scenarios. The voice agent transfers to a human when the caller's intent is not a status query, when confidence in the status readout is below 90%, or when the caller explicitly requests a human. The human queue is not reduced during the pilot; headcount reallocation is a separate business decision made after 90 days of stable operation.\"},\"name\":\"Does the voice agent replace human support agents?\"}]},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/us-insurer-voice-agent-shipment-status-4-week-pilot\/#breadcrumbs\",\"@type\":\"BreadcrumbList\",\"itemListElement\":[{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\",\"name\":\"Home\",\"position\":1},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/\",\"name\":\"Blog\",\"position\":2},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/us-insurer-voice-agent-shipment-status-4-week-pilot\/\",\"name\":\"How a 2,400-Person US Insurer Cut Shipment-Status Call Time by 67% in 4 Weeks\",\"position\":3}]},{\"@id\":\"https:\/\/blog.forfis.com#org\",\"@type\":\"Organization\",\"name\":\"Forfis\",\"url\":\"https:\/\/blog.forfis.com\"}]}","geo_content_hash":"ba4e0415c99dc51eecdab20879219b023f60b65bc77bde468e2ae88c09aefc8b","footnotes":""},"categories":[57],"tags":[67,73,23],"class_list":["post-483","post","type-post","status-publish","format-standard","hentry","category-insurance-and-insurtech","tag-order-and-shipment-status-updates","tag-replace-manual-data-entry","tag-usa"],"_links":{"self":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/483","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/comments?post=483"}],"version-history":[{"count":0,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/483\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/media?parent=483"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/categories?post=483"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/tags?post=483"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}