Pre-Pilot: Baseline and Infrastructure
-
Verify the contract volume and complexity profile. Count the number of MSAs, SOWs, and DPAs processed monthly by the legal team. This determines whether the pilot targets high-volume standard contracts or a narrower, higher-complexity subset. A B2B SaaS firm at 2,000+ employees typically processes 300-800 contracts per month across sales, procurement, and data-protection workflows.
-
Document the current review workflow end-to-end. Map each step from contract receipt to legal sign-off, including handoffs between paralegals, reviewers, and approvers. This baseline is the reference point for the before/after measurement. Without it, you cannot quantify cycle-time reduction or error-rate improvement after the pilot.
-
Define the standard playbook in Confluence. Consolidate the firm’s standard clauses, acceptable deviations, and red-flag categories into a structured Confluence space. The RAG pipeline retrieves from this space, so its completeness and clarity directly determine the agent’s accuracy. Ambiguous or outdated playbook entries will propagate into false positives.
-
Select the open-weight model and GPU infrastructure. Choose a model (e.g., Llama 3 70B or Mistral 8x7B) and provision on-premise GPU servers with at least 80 GB VRAM per node. On-premise deployment ensures no contract data leaves the building, which is a hard requirement for a compliance-safe rollout in Germany. The model must support English and German contract language.
-
Build the RAG index from historical contracts and playbook documents. Generate embeddings using a multilingual model (e.g., BGE-M3) and index all standard templates, reviewed contracts, and playbook entries. The index is the agent’s knowledge base. A poorly constructed index—missing key clause categories or containing outdated templates—will degrade retrieval quality and increase hallucination risk.
Pilot Build: Extraction, RAG, and Human-in-the-Loop
-
Configure the document extraction pipeline. Set up PDF and DOCX parsing to extract structured fields: parties, obligations, SLAs, termination clauses, and data-processing terms. The extraction pipeline feeds the RAG system and the classification model. Inaccurate extraction—missing a liability cap or misreading a termination date—will cascade into incorrect risk assessments. Test the pipeline on 50 historical contracts before proceeding.
-
Implement the human-in-the-loop approval workflow. Define which clause categories require mandatory human review (liability caps, data processing, termination rights) and configure the routing rules. The agent drafts and classifies, but a person approves anything that touches a contract. This is a policy constraint, not a model limitation. The workflow should enforce this via configuration, not rely on the model’s confidence score.
-
Set the error-rate targets and measurement protocol. Define the acceptable false-positive and false-negative rates (target: under 8% combined by month 6) and the cycle-time target (under 15 minutes for a standard 20-page MSA). These targets are the success criteria for the pilot. Without them, you cannot determine whether the system is ready for rollout or needs further tuning. The measurement protocol should specify how each metric is calculated and who is responsible for tracking it.
-
Deploy the pilot to a single contract type. Start with the highest-volume, lowest-complexity contract type—typically standard MSAs with a fixed clause set. This gives the model a clear training signal and a measurable baseline. Avoid starting with complex, multi-party agreements or contracts with significant negotiation history. The pilot should process at least 200 contracts to generate statistically meaningful error-rate data.
Pilot Execution: Feedback, Drift, and SOP
-
Run the pilot for 8 weeks with weekly feedback loops. Have the legal team review every agent-flagged clause and provide feedback on misclassifications. The feedback loop is the primary tuning mechanism. Without it, the model will not adapt to the firm’s specific contract language and risk appetite. Schedule a 30-minute weekly review with the legal team to discuss the top 10 misclassifications and adjust the playbook or prompts accordingly.
-
Monitor model drift and hallucination rates. Track the rate at which the agent generates clauses not present in the playbook or misattributes obligations to the wrong party. Hallucination is the primary risk in contract review. A single hallucinated liability clause can create legal exposure. Monitor this metric daily during the pilot and set an alert threshold at 2% hallucination rate. If the threshold is breached, pause the pilot and investigate the root cause.
-
Document the SOP for managed operations. Write a standard operating procedure covering model retraining frequency, RAG index update cadence, escalation paths, and audit-log retention. The SOP is the handover document for the managed operations phase. It should specify who is responsible for each task, how often it is performed, and what the acceptance criteria are. Without a documented SOP, the system will degrade as contract language evolves and the legal team’s risk appetite shifts.
Rollout and Managed Operations
-
Transition to managed operations with a defined SLA. Agree on the SLA for accuracy (under 8% combined error rate), cycle time (under 15 minutes), and availability (99.5% uptime). Managed operations means the vendor handles model retraining, prompt versioning, RAG index updates, and monitoring. The client’s legal team provides feedback, which feeds into a monthly retraining cycle. The SLA is the contractual basis for ongoing support and the trigger for remediation if performance degrades.
-
Establish the monthly performance reporting cadence. The vendor should provide a monthly report covering contracts processed, average cycle time, false-positive and false-negative rates, top 5 most-flagged clause categories, and model drift metrics. The legal team reviews this report and provides feedback on specific misclassifications. The vendor uses this feedback to retrain the model and update the RAG index. Quarterly, a joint review assesses whether the system meets the agreed SLA and whether scope expansion is justified.
-
Maintain the audit trail for compliance. Log every contract processed, the agent’s classification, the human reviewer’s decision, and the final outcome. This audit trail is stored in the client’s own infrastructure, not the vendor’s. Logs should be retained for at least 7 years to align with German commercial record-keeping requirements (HGB §257). The log format should be machine-readable (JSON) to support future compliance audits or regulatory inquiries.
-
Schedule quarterly scope reviews. Assess whether the system is ready to expand to additional contract types (DPAs, NDAs, procurement agreements) or jurisdictions. Scope expansion should be driven by the pilot’s performance data, not by ambition. If the combined error rate is consistently under 8% and the cycle-time target is met, the next contract type can be added to the RAG index and the pilot can be extended. If not, focus on tuning the current scope before expanding.