4-Week AI Contract Review Pilot for a 15-Person Swiss E-Commerce Team

The problem: contract review at 6.2 hours per document in a 15-person Swiss e-commerce team

A 15-person e-commerce and retail company in Switzerland reviews vendor onboarding agreements, customer return-policy acknowledgments, and marketplace seller terms by hand. Each contract takes a median of 6.2 hours from receipt to signed approval, and 11% of contracts ship with a missed clause or an incorrect term. The legal and compliance function is a single person who also handles GDPR inquiries and tax filings. The company needs multilingual coverage across English, German, and French, and it wants to lower the cost per support ticket without adding headcount. The constraint is a 4-week fixed-scope pilot: no open-ended discovery, no multi-department rollout in the first engagement. The deliverable is a measured before/after baseline on cycle time and error rate for one contract-review workflow, plus a 12-month scaling roadmap across departments.

Prerequisites before step 1

  • PostgreSQL 15 or later with the pgvector extension installed (CREATE EXTENSION vector;). The extension must be available on the client’s own instance; do not use a managed vector database for this pilot.
  • A contract library of at least 200 historical contracts in English, German, and French, exported as PDF or DOCX. These become the embedding index.
  • Google Workspace with API access enabled: the Drive API for document storage, the Gmail API for notifications, and the Chat API for approval workflows. The service account needs drive.file and gmail.send scopes.
  • An LLM API key for OpenAI (GPT-4o) or Anthropic (Claude 3.5 Sonnet). The key must have access to the text-embedding-3-small endpoint for the embedding step.
  • A single VM with 16 GB RAM and either an A10G GPU (24 GB VRAM) for batch embedding or a CPU-only setup if contract volume is under 500 per month.
  • One named reviewer from the legal and compliance function who will approve or reject every LLM-drafted clause during the pilot. This person must be available for 2 hours per day during weeks 3 and 4.

Step 1: Build the pgvector contract index

Export the 200 historical contracts from Google Drive to a local directory. Run a Python script that splits each contract into clauses using a regex on section headers (e.g., ^\d+\.\d+\s+[A-Z]). For each clause, call the text-embedding-3-small endpoint with the clause text and store the 1,536-dimensional vector in a contract_clauses table with columns id, contract_id, clause_text, embedding vector(1536), language, and created_at. The script should log the embedding latency per clause; expect 18 ms per call on a GPT-4o endpoint. After indexing, run a sanity check: embed a known clause and query the top-5 matches. If the original clause does not appear in the top-5, the index is broken and you must re-run the embedding step.

Step 2: Wire the workflow orchestration layer

Define the state machine in a YAML file with five states: received, embedded, drafted, awaiting_approval, and approved. The received state triggers the embedding step. The embedded state calls the LLM with the top-5 pgvector matches as context and the incoming contract clause as the query. The LLM returns a JSON object with suggested_revision, confidence_score, and flagged_terms. The drafted state sends a Google Chat message to the reviewer with the clause text, the suggested revision, and a link to the Google Doc. The awaiting_approval state pauses for 48 hours. If the reviewer approves, the state moves to approved and the contract is marked complete. If the reviewer rejects, the state returns to drafted with the reviewer’s comment appended to the LLM prompt. Log every state transition in a workflow_log table with the reviewer’s Google Workspace ID, the clause hash, and the timestamp.

Step 3: Run the human-in-the-loop review for 10 business days

Run the pilot on the highest-volume contract type: vendor onboarding agreements. For each incoming contract, the orchestration layer embeds the clauses, queries pgvector, and calls the LLM. The LLM drafts a revision for any clause that does not match the company’s standard template. The reviewer receives a Google Chat notification with the flagged clause and the suggested revision. The reviewer opens the contract in Google Docs, sees the flagged clause highlighted in yellow, and clicks approve or reject. The state machine records the decision. Run the pilot for 10 business days. Track three metrics per contract: cycle time (hours from receipt to approved), error rate (percentage of clauses the reviewer had to edit), and cost per ticket (LLM API cost + reviewer time × hourly rate). The baseline from the audit is 6.2 hours, 11% error rate, and CHF 42 per contract.

Step 4: Measure cycle time, error rate, and cost per ticket

At the end of the 10-day pilot, compare the measured metrics against the baseline. The go/no-go criteria are defined in the pilot contract: if cycle time drops by at least 50% (to 3.1 hours or less) and error rate drops by at least 40% (to 6.6% or less), the client proceeds to rollout. If either criterion is not met, the pilot is extended by 5 business days with a revised LLM prompt or a different embedding model. The measurement report includes a per-clause breakdown: which clause types the LLM handled well (e.g., payment terms, liability caps) and which still require human review (e.g., IP assignment, termination clauses). The report also includes the cost per ticket for the pilot period and a projection for 12 months at the current contract volume. The 12-month scaling roadmap identifies the next two workflows to automate: customer return-policy acknowledgments and marketplace seller terms.

Common pitfalls and how to detect them

  • Embedding drift: if the contract template changes (e.g., a new liability clause is added), the pgvector index becomes stale. Detect this by running a weekly job that embeds the current template and compares it against the index. If the top-5 match score drops below 0.82, re-index the affected clauses.
  • Reviewer bottleneck: if the reviewer does not respond within 48 hours, the workflow stalls. Detect this by monitoring the awaiting_approval state duration. If the median wait exceeds 36 hours, escalate to the team lead via a Gmail API email.
  • Language misclassification: if a German contract is misclassified as English, the LLM may produce a low-quality draft. Detect this by logging the detected language per contract and flagging any contract where the detected language does not match the contract’s metadata field.
  • LLM hallucination: if the LLM invents a clause that does not exist in the contract library, the reviewer will reject it. Detect this by logging the confidence_score and flagging any draft with a score below 0.70 for manual review before it reaches the reviewer.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *