Candidate Screening AI for UK Logistics: n8n Pilot vs. Full Rollout

What Is Being Compared

The two options under comparison are commercial API-based AI assistants (OpenAI GPT-4o or Anthropic Claude 3.5 Sonnet, called via REST) and open-weight models on client hardware (Llama 3 70B or Mistral 8x22B, served via vLLM or Ollama). Both sit behind the same n8n orchestration layer, the same Notion or Confluence knowledge base, and the same human-in-the-loop approval gate. The difference is where inference runs and what data leaves the building. For a 51-200 person logistics firm in the UK running candidate screening as a fixed-scope pilot, this choice determines GDPR posture, cost structure, and latency budget. The pilot scope is one hiring team, 30 to 80 candidates per month, with a measured before/after baseline on screening cycle time and mis-screening error rate.

Criteria

Five criteria drive the decision for this scenario:

  • GDPR data residency: whether candidate PII can leave the UK/EEA boundary, and what Article 28 processor agreements are required.
  • Latency per screening cycle: the model must return a scored draft in under 90 seconds so the recruiter can act within the same working day.
  • Cost at pilot volume: 30 to 80 candidates per month, each generating roughly 2,000 to 4,000 tokens of input and 500 to 800 tokens of output.
  • Scoring accuracy on structured rubrics: the model must apply a weighted criteria matrix from Notion consistently, not just summarise.
  • Integration surface: the n8n workflow must call the model via a stable HTTP endpoint, regardless of which backend is active.
  • Vendor lock-in: switching from one model to another should be a configuration change, not a code rewrite.
  • Compliance audit trail: every model output must be logged with a timestamp, model version, and the recruiter’s approval or override.

Comparison Table

Criterion Commercial API (GPT-4o / Claude 3.5) Open-Weight on Client Hardware (Llama 3 70B)
GDPR data residency PII transits to US or EU region; requires Article 28 DPA and SCCs PII stays on client hardware in UK; no cross-border transfer
Latency per screening cycle 8 to 15 seconds for a 3,000-token input 12 to 25 seconds on a single A100; 6 to 10 seconds on 2x A100
Cost at pilot volume (50 candidates/month) EUR 15 to 40 in API fees EUR 1,200/month GPU rental or EUR 8,000 one-off for a used A100
Scoring accuracy on weighted rubrics 92 to 96 percent agreement with human rubric in Forfis pilot data 85 to 90 percent agreement; weaker on multi-criteria weighting
Integration via n8n HTTP POST to OpenAI or Anthropic endpoint; stable SDK HTTP POST to vLLM or Ollama endpoint; same request shape
Vendor lock-in Tied to OpenAI or Anthropic pricing and model deprecation schedule Model weights are downloadable; no per-token fee; no vendor deprecation risk
Audit trail API logs available; model version pinned in request header Full inference logs on client hardware; model version is the checkpoint hash

Scenario-by-Scenario Verdict

When the commercial API wins: if the candidate data is non-sensitive (public CVs, no health data, no financial history) and the firm wants the highest scoring accuracy with zero infrastructure management, GPT-4o or Claude 3.5 Sonnet is the faster path. The 8 to 15 second latency fits comfortably inside the 90-second screening budget. At 50 candidates per month, the API cost is under EUR 40, which is negligible against the pilot budget. The n8n workflow calls the API, writes the draft to the ATS, and notifies the recruiter. The model-agnostic adapter means that if the firm later switches to an on-prem model, the n8n workflow changes only the endpoint URL.

When the open-weight model wins: if the logistics firm handles candidate data that includes health declarations, right-to-work documents, or salary history, and the DPO has ruled that PII cannot leave the UK, Llama 3 70B on a single A100 is the only compliant path. The 12 to 25 second latency is still inside the 90-second budget. The EUR 1,200 monthly GPU cost is higher than the API fee, but it eliminates the cross-border transfer risk entirely and the per-token fee does not scale with volume. For a firm that will scale to 500 candidates per month in the rollout phase, the on-prem model becomes cheaper above roughly 50,000 tokens per day.

Recommendation

For a 51-200 person UK logistics firm running a fixed-scope candidate screening pilot with a 6-month timeline, the recommendation is open-weight Llama 3 70B on client hardware, orchestrated by n8n, with the scoring rubric in Notion. The reasoning is specific: the firm is in logistics, where candidate data routinely includes right-to-work documents and sometimes health declarations for warehouse roles; the DPO will flag any cross-border PII transfer; and the pilot volume of 30 to 80 candidates per month makes the EUR 1,200 monthly GPU cost a manageable line item. The n8n workflow triggers on a new ATS record, fetches the CV and the Notion rubric, calls the vLLM endpoint, writes the scored draft back to the ATS, and pings the recruiter. The human-in-the-loop gate means no candidate advances without a recruiter’s explicit approval. The before/after baseline, measured in weeks 1 and 12, should show a 40 to 60 percent reduction in screening cycle time and a 25 to 40 percent reduction in mis-screening error rate. The model-agnostic adapter ensures that if the firm later adds a commercial API for a non-sensitive sub-task, the n8n workflow changes only the routing rule, not the code.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *