The Problem: Manual Contract Review in a 2,000+ Employee Logistics Firm
A 2,000+ employee logistics company in the USA processes hundreds of freight forwarding, warehouse, and vendor contracts monthly. Senior staff spend 3-5 hours per contract on manual clause review, with a 15-25% error rate on obligation identification. The cost per contract runs $250-400 in labor, and the cycle time delays onboarding by 5-10 business days. The problem is not a lack of tools but a lack of a structured pipeline that grounds LLM output in the company’s own policy documents and historical precedent while maintaining ISO 27001 audit trails. The pilot must reduce cycle time to under 90 minutes, cut error rates below 5%, and free senior staff for negotiation and exception work within 8 weeks.
Prerequisites Before Step 1
Before starting the pilot, confirm the following are in place:
- API access to the contract repository (e.g., DocuSign, iManage, or a shared drive) and the CRM (Salesforce, HubSpot) where contract metadata lives.
- Notion or Confluence workspace containing standard clause templates, internal policies, and approval workflows, with read API access enabled.
- PostgreSQL 15+ with the
pgvectorextension installed, provisioned on the client’s own infrastructure or a private cloud VPC to satisfy ISO 27001 data residency requirements. - LLM API keys for OpenAI (GPT-4o) or Anthropic (Claude 3.5 Sonnet) for the classification and drafting layer, with rate limits and cost caps configured.
- A named senior reviewer per contract type who will serve as the human-in-the-loop approver during the pilot.
- Baseline metrics documented: average cycle time, error rate, and cost per contract for the selected contract type over the last 90 days.
Step 1-3: Build the pgvector Retrieval Layer
-
Export and chunk policy documents. Pull all standard clause templates and policy statements from Notion or Confluence via their REST APIs. Chunk each document into 200-400 token segments with 50-token overlap. Store the raw text and chunk metadata (source URL, version, last-modified timestamp) in a
policy_chunkstable in PostgreSQL. -
Generate and store embeddings. Use the
text-embedding-3-smallmodel (OpenAI) ornomic-embed-text(open-weight, if data cannot leave the building) to generate 1536-dimensional vectors for each chunk. Insert them into apgvectortable with an HNSW index:CREATE INDEX ON policy_chunks USING hnsw (embedding vector_cosine_ops);. Verify index build time is under 5 minutes for 10k chunks. -
Build the retrieval function. Write a Python function that takes a contract clause string, embeds it, and queries
pgvectorfor the top-5 most similar policy chunks. Return the chunks with their cosine similarity scores. Set a minimum threshold of 0.75; below this, flag the clause for mandatory human review.
Step 4-6: LLM Classification and Human Approval
-
Integrate the LLM classification layer. For each extracted clause, construct a prompt that includes: (a) the clause text, (b) the top-5 retrieved policy chunks with their similarity scores, (c) the contract type and counterparty name. Instruct the model to classify the clause as
standard,modified, ornon-standard, and to extract all obligations with their source text spans. Use GPT-4o or Claude 3.5 Sonnet withtemperature=0.1for deterministic output. -
Add the human approval gate. Route every
modifiedornon-standardclause to the named senior reviewer via a simple web form or Slack integration. The reviewer sees the clause, the retrieved policy context, and the model’s classification. They approve, reject, or edit the classification. Log every decision with a timestamp and reviewer ID for ISO 27001 audit trails. -
Implement the secondary verification check. After the LLM extracts obligations, run a second LLM call that verifies each extracted obligation has a direct textual match in the source PDF. If the match score drops below 0.85, log a discrepancy and escalate to a senior reviewer. This catches hallucinated clauses before they reach the approval stage.
Step 7-9: Orchestration, UAT, and Handoff
-
Orchestrate the workflow with state tracking. Use Temporal, n8n, or a custom Python state machine to track each contract through stages:
ingested,clauses_extracted,classified,pending_approval,approved,signed. Each stage has a timeout (30 minutes for extraction, 4 hours for approval) and a fallback action (escalate to a senior reviewer if approval is not received). Log every state transition with a timestamp, actor, and input/output hashes. Store logs in an append-only table to satisfy ISO 27001 audit requirements. -
Run UAT with 20-30 real contracts. Select a mix of standard and complex contracts from the last 90 days. Measure cycle time, error rate, and cost per contract. Compare against the baseline. Target: cycle time under 90 minutes, error rate under 5%, cost per contract under $30. Document all discrepancies and feed them back into the prompt and retrieval thresholds.
-
Collect ISO 27001 evidence and hand off. Export the audit logs, access control records, and data retention policies. Document the system architecture, API call logs, and encryption configurations. Hand off to the operations team with a runbook covering model version updates, pgvector index maintenance, and escalation paths. The next logical step is to expand the pilot to a second contract type and integrate with the ERP for automated PO generation.
Common Pitfalls and How to Detect Them
-
Hallucinated clauses. The model invents obligations not present in the source document. Detect via the secondary verification check (match score below 0.85) and the retrieval confidence threshold (below 0.75). Without these guardrails, a single hallucinated indemnity clause can create a $2M+ liability exposure.
-
Stale policy context. The pgvector index contains outdated clause templates because the Notion/Confluence sync failed. Detect by checking the
last_syncedtimestamp in thepolicy_chunkstable and alerting if it exceeds 24 hours. Run a nightly sync job and log failures. -
Approval bottleneck. Senior reviewers do not respond within the 4-hour window, stalling the pipeline. Detect by monitoring the
pending_approvalstate duration. Escalate to a backup reviewer after 2 hours and log the escalation for process improvement. -
API cost overrun. Unbounded LLM calls on large contracts (50+ pages) drive API costs above budget. Detect by logging token counts per call and setting a hard cap of 50k tokens per contract. Chunk large contracts and process them in batches.
-
ISO 27001 audit gap. Missing logs for API calls or access control changes. Detect by running a weekly audit log integrity check that verifies every state transition has a corresponding log entry with a hash. Alert on any gaps.
Leave a Reply