Cutting Contract First-Response Time to 4 Hours: A Swiss E-Commerce AI Pilot

Background: A Zurich E-Commerce Firm at 340 Heads

This case study is a composite drawn from patterns Forfis has observed across multiple engagements in Tier-1 European markets. No named customer is represented; the company, metrics, and timeline are representative of the median engagement in this segment.

The company is a mid-market e-commerce and retail operator based in Zurich, with roughly 340 employees across operations, logistics, and customer service. It runs a B2B2C model: wholesale contracts with 120+ regional retailers, plus direct-to-consumer sales through its own web platform. The legal and compliance team consists of six in-house lawyers and two external counsel retained for high-value or cross-border deals. The existing stack includes SAP S/4HANA for ERP, Salesforce for CRM, and Microsoft 365 with Teams as the primary collaboration layer. Contract documents arrive as PDFs and Word files through email and a shared SharePoint drive, and every one of them passes through a manual review queue before the legal team signs off.

The company is in the scaling phase of its AI adoption: it had piloted a basic document classification model in 2023 but had not yet extended AI tooling beyond a single department. The legal team was the next logical target, given the volume of incoming contracts and the recurring nature of the review work.

Challenge: 48-Hour First-Response Time and a Flat Headcount

The legal team was processing an average of 45 to 60 contracts per week across wholesale agreements, retailer onboarding documents, and supplier terms. The median first-response time — the interval from contract receipt to the first substantive legal annotation — was 48 hours. For high-value contracts exceeding CHF 250,000, the figure stretched to 72 hours or more. The bottleneck was not the lawyers’ expertise but the triage step: a junior associate had to read every incoming document, classify its type, flag non-standard clauses, and route it to the appropriate senior reviewer before any substantive work began.

Three pressures made the status quo unsustainable. First, the company was onboarding 15 to 20 new regional retailers per quarter, each requiring a customized wholesale agreement with variable payment terms, return policies, and liability caps. Second, the EU AI Act’s phased application timeline meant that any AI system deployed for contract review would need to meet Article 50 transparency and Article 14 human-oversight requirements by August 2026, and the legal team wanted the compliance documentation built into the tool from the start rather than retrofitted. Third, headcount was flat: the company had no budget to add a seventh lawyer, and the external counsel retainer was already at CHF 18,000 per month.

The operational target was explicit: cut first-response time to under 6 hours for standard contracts and under 24 hours for high-value ones, without increasing legal headcount.

Approach: pgvector Retrieval, Predictive Scoring, and a Teams Integration

Forfis engaged as a dedicated AI team of four: a technical lead, a product designer, a full-stack engineer, and a domain specialist with legal-tech experience. The engagement ran over six months, structured as a fixed-scope pilot on the contract review workflow before any rollout to other departments.

The architecture was model-agnostic by design. For clause classification and risk scoring, the system used OpenAI’s GPT-4o API, which handled the nuanced language of Swiss commercial law with acceptable accuracy on the pilot’s evaluation set. For the retrieval layer, the team built a pgvector index in PostgreSQL, storing embeddings of the company’s 2,400 historical contracts, 380 internal policy documents, and the relevant Swiss Code of Obligations (OR) articles. Each incoming contract was chunked into clause-level segments, embedded using text-embedding-3-small (1,536 dimensions), and matched against the index via cosine similarity. The top 8 retrieved passages were injected into the LLM’s context window, grounding its output in the company’s own precedent rather than general training data.

The predictive scoring model assigned a 0-100 risk score to each contract based on clause deviation, non-standard liability language, and historical dispute frequency. Contracts scoring above 75 routed to mandatory human review; those below 40 auto-approved for standard terms. The middle band (40-75) received AI-drafted annotations but required a human sign-off. Every decision was logged with a timestamp, the model version, and the retrieved context, satisfying the EU AI Act’s audit-trail requirements under Article 12.

The integration point was Microsoft Teams. When a contract was uploaded to the SharePoint drive, a Power Automate flow triggered the AI pipeline, and the resulting risk score, clause annotations, and suggested redlines appeared as a card in the legal team’s designated Teams channel. The reviewer approved or rejected with a single click, and the decision was written back to Salesforce and the SharePoint metadata.

Outcome: 4.2-Hour First-Response and a 3.1% Residual Error Rate

The pilot ran for eight weeks after the build phase, covering approximately 380 contracts across the three categories. The measured outcomes, compared against the pre-pilot baseline:

  • First-response time for standard contracts dropped from a median of 48 hours to 4.2 hours. For high-value contracts, the median fell from 72 hours to 19 hours. The reduction came primarily from eliminating the manual triage step; the AI classified and scored the contract within 90 seconds of upload, and the Teams notification reached the reviewer in under 2 minutes.

  • Error rate on clause classification (measured as the percentage of clauses misclassified by the AI versus the legal team’s final determination) was 6.8% in the first two weeks of the pilot and stabilized at 3.1% by week eight after prompt refinement and threshold adjustment. The human-in-the-loop gate caught every misclassification before it reached a signed contract.

  • Reviewer throughput increased: the same six lawyers processed 58 contracts per week during the pilot versus 45 in the baseline period, a 29% increase without additional headcount.

  • External counsel spend on routine contract review fell by an estimated 35%, as the AI handled the first-pass annotation for standard terms, leaving external counsel engaged only on genuinely novel or cross-border issues.

The EU AI Act compliance file — including the model’s intended purpose statement, the human-oversight protocol, the data governance log, and the evaluation metrics — was delivered as a standalone document in week 22, ahead of the August 2026 high-risk system deadline.

Lessons for Teams Scaling AI Across Departments

  • Baseline before you build. The 48-hour median and the 6.8% initial error rate were only meaningful because the team measured them before writing a line of code. Without the pre-pilot baseline, the 4.2-hour outcome would have been an anecdote rather than a defensible metric. Every pilot in this segment should ship with a measured before/after on cycle time and error rate, not a qualitative “faster” claim.

  • Retrieval quality determines ceiling. The pgvector index was the single highest-leverage component. When the team expanded the index from 2,400 to 4,100 documents (adding two years of archived contracts and the full OR text), the classification error rate dropped from 3.1% to 2.4% without any change to the LLM or the prompt. Teams scaling across departments should treat the retrieval corpus as a first-class asset, not an afterthought.

  • Human-in-the-loop is not a safety net; it is the product. The approval gate in Teams was where the legal team’s domain knowledge fed back into the system. Every rejection with a comment became a training signal for the next prompt iteration. Removing the human gate to “speed things up” would have eliminated the feedback loop that kept the error rate below 4%.

  • Compliance is a build-time constraint, not a launch-time checkbox. The EU AI Act documentation was produced in week 22, not week 24. Building the audit log, the model versioning, and the human-oversight protocol into the architecture from week 5 meant the compliance file was a documentation exercise, not a re-engineering project. Teams facing the August 2026 deadline should start the compliance file in the first sprint, not the last.

  • Model-agnosticism is an operational hedge, not a theoretical preference. When OpenAI’s API pricing changed in month 4, the team rerouted 40% of the classification volume to an on-premises Llama 3 70B instance for the lower-complexity contract types, reducing API spend by 22% without degrading accuracy below the 3.1% threshold. The abstraction layer made this a configuration change, not a re-architecture.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *