{"id":502,"date":"2026-10-06T19:00:46","date_gmt":"2026-10-06T19:00:46","guid":{"rendered":"https:\/\/blog.forfis.com\/blog\/ai-contract-review-pilot-swiss-ecommerce-pgvector\/"},"modified":"2026-10-06T19:00:46","modified_gmt":"2026-10-06T19:00:46","slug":"ai-contract-review-pilot-swiss-ecommerce-pgvector","status":"publish","type":"post","link":"https:\/\/blog.forfis.com\/blog\/ai-contract-review-pilot-swiss-ecommerce-pgvector\/","title":{"rendered":"4-Week AI Contract Review Pilot for a 15-Person Swiss E-Commerce Team"},"content":{"rendered":"<h2>The problem: contract review at 6.2 hours per document in a 15-person Swiss e-commerce team<\/h2>\n<p>A 15-person e-commerce and retail company in Switzerland reviews vendor onboarding agreements, customer return-policy acknowledgments, and marketplace seller terms by hand. Each contract takes a median of 6.2 hours from receipt to signed approval, and 11% of contracts ship with a missed clause or an incorrect term. The legal and compliance function is a single person who also handles GDPR inquiries and tax filings. The company needs multilingual coverage across English, German, and French, and it wants to lower the cost per support ticket without adding headcount. The constraint is a 4-week fixed-scope pilot: no open-ended discovery, no multi-department rollout in the first engagement. The deliverable is a measured before\/after baseline on cycle time and error rate for one contract-review workflow, plus a 12-month scaling roadmap across departments.<\/p>\n<h2>Prerequisites before step 1<\/h2>\n<ul>\n<li><strong>PostgreSQL 15 or later<\/strong> with the <code>pgvector<\/code> extension installed (<code>CREATE EXTENSION vector;<\/code>). The extension must be available on the client\u2019s own instance; do not use a managed vector database for this pilot.<\/li>\n<li><strong>A contract library<\/strong> of at least 200 historical contracts in English, German, and French, exported as PDF or DOCX. These become the embedding index.<\/li>\n<li><strong>Google Workspace<\/strong> with API access enabled: the Drive API for document storage, the Gmail API for notifications, and the Chat API for approval workflows. The service account needs <code>drive.file<\/code> and <code>gmail.send<\/code> scopes.<\/li>\n<li><strong>An LLM API key<\/strong> for OpenAI (GPT-4o) or Anthropic (Claude 3.5 Sonnet). The key must have access to the <code>text-embedding-3-small<\/code> endpoint for the embedding step.<\/li>\n<li><strong>A single VM<\/strong> with 16 GB RAM and either an A10G GPU (24 GB VRAM) for batch embedding or a CPU-only setup if contract volume is under 500 per month.<\/li>\n<li><strong>One named reviewer<\/strong> from the legal and compliance function who will approve or reject every LLM-drafted clause during the pilot. This person must be available for 2 hours per day during weeks 3 and 4.<\/li>\n<\/ul>\n<h2>Step 1: Build the pgvector contract index<\/h2>\n<p>Export the 200 historical contracts from Google Drive to a local directory. Run a Python script that splits each contract into clauses using a regex on section headers (e.g., <code>^\\d+\\.\\d+\\s+[A-Z]<\/code>). For each clause, call the <code>text-embedding-3-small<\/code> endpoint with the clause text and store the 1,536-dimensional vector in a <code>contract_clauses<\/code> table with columns <code>id<\/code>, <code>contract_id<\/code>, <code>clause_text<\/code>, <code>embedding vector(1536)<\/code>, <code>language<\/code>, and <code>created_at<\/code>. The script should log the embedding latency per clause; expect 18 ms per call on a GPT-4o endpoint. After indexing, run a sanity check: embed a known clause and query the top-5 matches. If the original clause does not appear in the top-5, the index is broken and you must re-run the embedding step.<\/p>\n<h2>Step 2: Wire the workflow orchestration layer<\/h2>\n<p>Define the state machine in a YAML file with five states: <code>received<\/code>, <code>embedded<\/code>, <code>drafted<\/code>, <code>awaiting_approval<\/code>, and <code>approved<\/code>. The <code>received<\/code> state triggers the embedding step. The <code>embedded<\/code> state calls the LLM with the top-5 pgvector matches as context and the incoming contract clause as the query. The LLM returns a JSON object with <code>suggested_revision<\/code>, <code>confidence_score<\/code>, and <code>flagged_terms<\/code>. The <code>drafted<\/code> state sends a Google Chat message to the reviewer with the clause text, the suggested revision, and a link to the Google Doc. The <code>awaiting_approval<\/code> state pauses for 48 hours. If the reviewer approves, the state moves to <code>approved<\/code> and the contract is marked complete. If the reviewer rejects, the state returns to <code>drafted<\/code> with the reviewer\u2019s comment appended to the LLM prompt. Log every state transition in a <code>workflow_log<\/code> table with the reviewer\u2019s Google Workspace ID, the clause hash, and the timestamp.<\/p>\n<h2>Step 3: Run the human-in-the-loop review for 10 business days<\/h2>\n<p>Run the pilot on the highest-volume contract type: vendor onboarding agreements. For each incoming contract, the orchestration layer embeds the clauses, queries pgvector, and calls the LLM. The LLM drafts a revision for any clause that does not match the company\u2019s standard template. The reviewer receives a Google Chat notification with the flagged clause and the suggested revision. The reviewer opens the contract in Google Docs, sees the flagged clause highlighted in yellow, and clicks approve or reject. The state machine records the decision. Run the pilot for 10 business days. Track three metrics per contract: cycle time (hours from receipt to approved), error rate (percentage of clauses the reviewer had to edit), and cost per ticket (LLM API cost + reviewer time \u00d7 hourly rate). The baseline from the audit is 6.2 hours, 11% error rate, and CHF 42 per contract.<\/p>\n<h2>Step 4: Measure cycle time, error rate, and cost per ticket<\/h2>\n<p>At the end of the 10-day pilot, compare the measured metrics against the baseline. The go\/no-go criteria are defined in the pilot contract: if cycle time drops by at least 50% (to 3.1 hours or less) and error rate drops by at least 40% (to 6.6% or less), the client proceeds to rollout. If either criterion is not met, the pilot is extended by 5 business days with a revised LLM prompt or a different embedding model. The measurement report includes a per-clause breakdown: which clause types the LLM handled well (e.g., payment terms, liability caps) and which still require human review (e.g., IP assignment, termination clauses). The report also includes the cost per ticket for the pilot period and a projection for 12 months at the current contract volume. The 12-month scaling roadmap identifies the next two workflows to automate: customer return-policy acknowledgments and marketplace seller terms.<\/p>\n<h2>Common pitfalls and how to detect them<\/h2>\n<ul>\n<li><strong>Embedding drift<\/strong>: if the contract template changes (e.g., a new liability clause is added), the pgvector index becomes stale. Detect this by running a weekly job that embeds the current template and compares it against the index. If the top-5 match score drops below 0.82, re-index the affected clauses.<\/li>\n<li><strong>Reviewer bottleneck<\/strong>: if the reviewer does not respond within 48 hours, the workflow stalls. Detect this by monitoring the <code>awaiting_approval<\/code> state duration. If the median wait exceeds 36 hours, escalate to the team lead via a Gmail API email.<\/li>\n<li><strong>Language misclassification<\/strong>: if a German contract is misclassified as English, the LLM may produce a low-quality draft. Detect this by logging the detected language per contract and flagging any contract where the detected language does not match the contract\u2019s metadata field.<\/li>\n<li><strong>LLM hallucination<\/strong>: if the LLM invents a clause that does not exist in the contract library, the reviewer will reject it. Detect this by logging the <code>confidence_score<\/code> and flagging any draft with a score below 0.70 for manual review before it reaches the reviewer.<\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>A 4-week fixed-scope pilot for a 15-person Swiss e-commerce team: audit, pgvector contract index, workflow orchestration, and measured cycle-time reduction.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"4-Week AI Contract Review Pilot for a 15-Person Swiss E-Commerce Team","rank_math_description":"A 4-week fixed-scope pilot for a 15-person Swiss e-commerce team: audit, pgvector contract index, workflow orchestration, and measured cycle-time reduction.","rank_math_focus_keyword":"multilingual support coverage contract review","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","_yoast_wpseo_focuskw":"","pll_lang":"en","geo_jsonld":"{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-contract-review-pilot-swiss-ecommerce-pgvector\/#article\",\"@type\":\"Article\",\"author\":{\"@id\":\"https:\/\/blog.forfis.com#org\"},\"dateModified\":\"2026-10-06T00:03:51.126138637+00:00\",\"datePublished\":\"2026-10-06T00:03:51.126138637+00:00\",\"description\":\"A 4-week fixed-scope pilot for a 15-person Swiss e-commerce team: audit, pgvector contract index, workflow orchestration, and measured cycle-time reduction.\",\"headline\":\"4-Week AI Contract Review Pilot for a 15-Person Swiss E-Commerce Team\",\"inLanguage\":\"en\",\"keywords\":[\"Scaling Across Departments\",\"pgvector Embeddings Search\",\"Workflow Orchestration\",\"Legal and Compliance\",\"11-50\",\"None\",\"Fixed-Scope Pilot\",\"E-commerce and Retail\",\"Google Workspace\",\"English\",\"Multilingual Support Coverage\",\"Switzerland\",\"4 weeks\",\"Contract Review\"],\"mainEntityOfPage\":\"https:\/\/blog.forfis.com\/blog\/ai-contract-review-pilot-swiss-ecommerce-pgvector\/\",\"publisher\":{\"@id\":\"https:\/\/blog.forfis.com#org\"}},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-contract-review-pilot-swiss-ecommerce-pgvector\/#faq\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"A 4-week fixed-scope pilot in a 15-person Swiss e-commerce team typically covers one contract-review workflow end-to-end. Week 1 is the process audit and baseline measurement. Week 2 builds the pgvector index and the orchestration layer. Week 3 runs the human-in-the-loop review with your legal staff. Week 4 measures cycle time and error rate against the baseline and documents the go\/no-go criteria for rollout. The pilot does not include multi-department scaling, which is a separate engagement.\"},\"name\":\"What does a 4-week fixed-scope pilot for contract review actually deliver?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"pgvector is a PostgreSQL extension that stores vector embeddings and performs similarity search natively inside your existing database. For contract review, you embed each clause or section of your contract library into 1,536-dimensional vectors (for OpenAI's text-embedding-3-small) or 768-dimensional vectors (for a local model like BGE-M3). At review time, the incoming contract's clauses are embedded and matched against the index. The top-k results feed the LLM prompt as context. This keeps your proprietary contract language on your own infrastructure and avoids sending raw contract text to a third-party vector database.\"},\"name\":\"What is pgvector and why use it for contract review instead of a dedicated vector database?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"For a Swiss e-commerce company handling contracts in English, German, and French, you need a multilingual embedding model. BGE-M3 (BAAI General Embedding, Multilingual) supports 100+ languages and produces 1,024-dimensional vectors. It runs on a single A10G GPU (24 GB VRAM) or on CPU with acceptable latency for batch indexing. For the LLM layer, use a model with strong multilingual instruction-following: GPT-4o or Claude 3.5 Sonnet handle English, German, and French contract language with consistent quality. The orchestration layer routes each contract to the appropriate language pipeline and logs the language detected per document.\"},\"name\":\"How do you handle multilingual contract review in English, German, and French?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The audit identifies 3-5 contract-review workflows worth automating. For a 15-person e-commerce team, the typical candidates are: vendor onboarding agreements, customer return\/refund policy acknowledgments, and marketplace seller terms. The audit measures baseline cycle time (median hours from receipt to signed approval) and error rate (percentage of contracts with a missed clause or incorrect term). The roadmap prioritizes by volume \u00d7 cycle time \u00d7 error rate. The fixed-scope pilot targets the highest-scoring workflow. The roadmap document includes a 12-month scaling plan across departments with quarterly milestones.\"},\"name\":\"What does the AI process audit produce for a 15-person e-commerce company?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The orchestration layer uses a state machine with explicit approval gates. When the LLM flags a clause as non-standard, the workflow pauses and sends a notification to the assigned reviewer via Google Workspace (email or Google Chat). The reviewer opens the contract in Google Docs, sees the flagged clause highlighted with the LLM's suggested revision, and either approves, edits, or rejects. The state machine records the decision and timestamp. If the reviewer does not respond within 48 hours, an escalation email goes to the team lead. The entire interaction is logged in a SQLite or PostgreSQL audit table with the reviewer's Google Workspace ID, the clause hash, and the decision.\"},\"name\":\"How does the human-in-the-loop approval work in the workflow orchestration?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot ships with a before\/after measurement report. Before: median cycle time per contract (e.g., 6.2 hours), error rate (e.g., 11% of contracts had a missed clause), and cost per ticket (e.g., CHF 42 per contract review). After: median cycle time (e.g., 1.8 hours), error rate (e.g., 3%), and cost per ticket (e.g., CHF 14). The report includes a per-clause breakdown showing which clause types the LLM handled well and which still require human review. The go\/no-go criteria are defined in the pilot contract: if cycle time drops by at least 50% and error rate drops by at least 40%, the client proceeds to rollout.\"},\"name\":\"What metrics does the pilot measure and what are the go\/no-go criteria?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The orchestration layer uses the Google Workspace API to send notifications and collect approvals. The LLM runs on OpenAI or Anthropic APIs for the drafting and classification steps. The pgvector index runs on the client's own PostgreSQL instance (version 15 or later with the pgvector extension). The contract documents are stored in Google Drive, accessed via the Drive API. The entire stack runs on the client's existing infrastructure: a single VM with 16 GB RAM and an A10G GPU for embedding, or a CPU-only setup if the contract volume is under 500 per month. No new SaaS subscriptions are required beyond the LLM API calls.\"},\"name\":\"What infrastructure is needed to run this stack in-house?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The orchestration layer is model-agnostic by design. The LLM call is abstracted behind an interface that accepts a model name and parameters. For the pilot, you use GPT-4o or Claude 3.5 Sonnet for quality. For rollout, if the client's legal team requires that contract text never leaves the building, you swap to a local model: Llama 3.1 70B or Mistral Large on a single A100 GPU. The pgvector index and orchestration layer remain unchanged. The only configuration change is the model endpoint URL and the API key. The human-in-the-loop approval gates remain identical regardless of which model generates the draft.\"},\"name\":\"Can the LLM be swapped to a local model without changing the orchestration layer?\"}]},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-contract-review-pilot-swiss-ecommerce-pgvector\/#breadcrumbs\",\"@type\":\"BreadcrumbList\",\"itemListElement\":[{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\",\"name\":\"Home\",\"position\":1},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/\",\"name\":\"Blog\",\"position\":2},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/ai-contract-review-pilot-swiss-ecommerce-pgvector\/\",\"name\":\"4-Week AI Contract Review Pilot for a 15-Person Swiss E-Commerce Team\",\"position\":3}]},{\"@id\":\"https:\/\/blog.forfis.com#org\",\"@type\":\"Organization\",\"name\":\"Forfis\",\"url\":\"https:\/\/blog.forfis.com\"}]}","geo_content_hash":"4e126c2874f2520dd9fe7c68f1cbaea0b3930d82e9dda487ace430d4394d523f","footnotes":""},"categories":[65],"tags":[31,33,43],"class_list":["post-502","post","type-post","status-publish","format-standard","hentry","category-e-commerce-and-retail","tag-contract-review","tag-multilingual-support-coverage","tag-switzerland"],"_links":{"self":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/502","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/comments?post=502"}],"version-history":[{"count":0,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/502\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/media?parent=502"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/categories?post=502"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/tags?post=502"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}