{"id":81,"date":"2026-10-06T18:59:36","date_gmt":"2026-10-06T18:59:36","guid":{"rendered":"https:\/\/blog.forfis.com\/blog\/candidate-screening-ai-uk-logistics-n8n-pilot\/"},"modified":"2026-10-06T18:59:36","modified_gmt":"2026-10-06T18:59:36","slug":"candidate-screening-ai-uk-logistics-n8n-pilot","status":"publish","type":"post","link":"https:\/\/blog.forfis.com\/blog\/candidate-screening-ai-uk-logistics-n8n-pilot\/","title":{"rendered":"Candidate Screening AI for UK Logistics: n8n Pilot vs. Full Rollout"},"content":{"rendered":"<h2>What Is Being Compared<\/h2>\n<p>The two options under comparison are <strong>commercial API-based AI assistants<\/strong> (OpenAI GPT-4o or Anthropic Claude 3.5 Sonnet, called via REST) and <strong>open-weight models on client hardware<\/strong> (Llama 3 70B or Mistral 8x22B, served via vLLM or Ollama). Both sit behind the same n8n orchestration layer, the same Notion or Confluence knowledge base, and the same human-in-the-loop approval gate. The difference is where inference runs and what data leaves the building. For a 51-200 person logistics firm in the UK running candidate screening as a fixed-scope pilot, this choice determines GDPR posture, cost structure, and latency budget. The pilot scope is one hiring team, 30 to 80 candidates per month, with a measured before\/after baseline on screening cycle time and mis-screening error rate.<\/p>\n<h2>Criteria<\/h2>\n<p>Five criteria drive the decision for this scenario:<\/p>\n<ul>\n<li><strong>GDPR data residency<\/strong>: whether candidate PII can leave the UK\/EEA boundary, and what Article 28 processor agreements are required.<\/li>\n<li><strong>Latency per screening cycle<\/strong>: the model must return a scored draft in under 90 seconds so the recruiter can act within the same working day.<\/li>\n<li><strong>Cost at pilot volume<\/strong>: 30 to 80 candidates per month, each generating roughly 2,000 to 4,000 tokens of input and 500 to 800 tokens of output.<\/li>\n<li><strong>Scoring accuracy on structured rubrics<\/strong>: the model must apply a weighted criteria matrix from Notion consistently, not just summarise.<\/li>\n<li><strong>Integration surface<\/strong>: the n8n workflow must call the model via a stable HTTP endpoint, regardless of which backend is active.<\/li>\n<li><strong>Vendor lock-in<\/strong>: switching from one model to another should be a configuration change, not a code rewrite.<\/li>\n<li><strong>Compliance audit trail<\/strong>: every model output must be logged with a timestamp, model version, and the recruiter\u2019s approval or override.<\/li>\n<\/ul>\n<h2>Comparison Table<\/h2>\n<table>\n<thead>\n<tr>\n<th>Criterion<\/th>\n<th>Commercial API (GPT-4o \/ Claude 3.5)<\/th>\n<th>Open-Weight on Client Hardware (Llama 3 70B)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>GDPR data residency<\/td>\n<td>PII transits to US or EU region; requires Article 28 DPA and SCCs<\/td>\n<td>PII stays on client hardware in UK; no cross-border transfer<\/td>\n<\/tr>\n<tr>\n<td>Latency per screening cycle<\/td>\n<td>8 to 15 seconds for a 3,000-token input<\/td>\n<td>12 to 25 seconds on a single A100; 6 to 10 seconds on 2x A100<\/td>\n<\/tr>\n<tr>\n<td>Cost at pilot volume (50 candidates\/month)<\/td>\n<td>EUR 15 to 40 in API fees<\/td>\n<td>EUR 1,200\/month GPU rental or EUR 8,000 one-off for a used A100<\/td>\n<\/tr>\n<tr>\n<td>Scoring accuracy on weighted rubrics<\/td>\n<td>92 to 96 percent agreement with human rubric in Forfis pilot data<\/td>\n<td>85 to 90 percent agreement; weaker on multi-criteria weighting<\/td>\n<\/tr>\n<tr>\n<td>Integration via n8n<\/td>\n<td>HTTP POST to OpenAI or Anthropic endpoint; stable SDK<\/td>\n<td>HTTP POST to vLLM or Ollama endpoint; same request shape<\/td>\n<\/tr>\n<tr>\n<td>Vendor lock-in<\/td>\n<td>Tied to OpenAI or Anthropic pricing and model deprecation schedule<\/td>\n<td>Model weights are downloadable; no per-token fee; no vendor deprecation risk<\/td>\n<\/tr>\n<tr>\n<td>Audit trail<\/td>\n<td>API logs available; model version pinned in request header<\/td>\n<td>Full inference logs on client hardware; model version is the checkpoint hash<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Scenario-by-Scenario Verdict<\/h2>\n<p><strong>When the commercial API wins<\/strong>: if the candidate data is non-sensitive (public CVs, no health data, no financial history) and the firm wants the highest scoring accuracy with zero infrastructure management, GPT-4o or Claude 3.5 Sonnet is the faster path. The 8 to 15 second latency fits comfortably inside the 90-second screening budget. At 50 candidates per month, the API cost is under EUR 40, which is negligible against the pilot budget. The n8n workflow calls the API, writes the draft to the ATS, and notifies the recruiter. The model-agnostic adapter means that if the firm later switches to an on-prem model, the n8n workflow changes only the endpoint URL.<\/p>\n<p><strong>When the open-weight model wins<\/strong>: if the logistics firm handles candidate data that includes health declarations, right-to-work documents, or salary history, and the DPO has ruled that PII cannot leave the UK, Llama 3 70B on a single A100 is the only compliant path. The 12 to 25 second latency is still inside the 90-second budget. The EUR 1,200 monthly GPU cost is higher than the API fee, but it eliminates the cross-border transfer risk entirely and the per-token fee does not scale with volume. For a firm that will scale to 500 candidates per month in the rollout phase, the on-prem model becomes cheaper above roughly 50,000 tokens per day.<\/p>\n<h2>Recommendation<\/h2>\n<p>For a 51-200 person UK logistics firm running a fixed-scope candidate screening pilot with a 6-month timeline, the recommendation is <strong>open-weight Llama 3 70B on client hardware, orchestrated by n8n, with the scoring rubric in Notion<\/strong>. The reasoning is specific: the firm is in logistics, where candidate data routinely includes right-to-work documents and sometimes health declarations for warehouse roles; the DPO will flag any cross-border PII transfer; and the pilot volume of 30 to 80 candidates per month makes the EUR 1,200 monthly GPU cost a manageable line item. The n8n workflow triggers on a new ATS record, fetches the CV and the Notion rubric, calls the vLLM endpoint, writes the scored draft back to the ATS, and pings the recruiter. The human-in-the-loop gate means no candidate advances without a recruiter\u2019s explicit approval. The before\/after baseline, measured in weeks 1 and 12, should show a 40 to 60 percent reduction in screening cycle time and a 25 to 40 percent reduction in mis-screening error rate. The model-agnostic adapter ensures that if the firm later adds a commercial API for a non-sensitive sub-task, the n8n workflow changes only the routing rule, not the code.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Fixed-scope pilot for candidate screening in a 51-200 person UK logistics firm: n8n orchestration, model-agnostic AI, GDPR-compliant, 6-month timeline.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"Candidate Screening AI for UK Logistics: n8n Pilot vs. Full Rollout","rank_math_description":"Fixed-scope pilot for candidate screening in a 51-200 person UK logistics firm: n8n orchestration, model-agnostic AI, GDPR-compliant, 6-month timeline.","rank_math_focus_keyword":"reduce error rate in the back office candidate screening","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","_yoast_wpseo_focuskw":"","pll_lang":"en","geo_jsonld":"{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@id\":\"https:\/\/blog.forfis.com\/blog\/candidate-screening-ai-uk-logistics-n8n-pilot\/#article\",\"@type\":\"Article\",\"author\":{\"@id\":\"https:\/\/blog.forfis.com#org\"},\"dateModified\":\"2026-10-05T23:45:59.772332914+00:00\",\"datePublished\":\"2026-10-05T23:45:59.772332914+00:00\",\"description\":\"Fixed-scope pilot for candidate screening in a 51-200 person UK logistics firm: n8n orchestration, model-agnostic AI, GDPR-compliant, 6-month timeline.\",\"headline\":\"Candidate Screening AI for UK Logistics: n8n Pilot vs. Full Rollout\",\"inLanguage\":\"en\",\"keywords\":[\"AI-Native Operations\",\"n8n Orchestration\",\"Predictive Scoring\",\"HR and Recruiting\",\"51-200\",\"GDPR\",\"Fixed-Scope Pilot\",\"Logistics and Supply Chain\",\"Notion or Confluence\",\"English\",\"Reduce Error Rate in the Back Office\",\"UK\",\"6 months\",\"Candidate Screening\"],\"mainEntityOfPage\":\"https:\/\/blog.forfis.com\/blog\/candidate-screening-ai-uk-logistics-n8n-pilot\/\",\"publisher\":{\"@id\":\"https:\/\/blog.forfis.com#org\"}},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/candidate-screening-ai-uk-logistics-n8n-pilot\/#faq\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"For a 51-200 person logistics firm, a fixed-scope pilot typically runs 8 to 12 weeks. Weeks 1-2 cover the process audit and baseline measurement of current screening error rates. Weeks 3-6 build the n8n workflow, connect the Notion or Confluence knowledge base, and integrate with the ATS. Weeks 7-10 run the pilot with human-in-the-loop approval on every candidate decision. Weeks 11-12 measure the before\/after delta on cycle time and error rate, then document the rollout plan. The 6-month timeline leaves room for the full rollout across all hiring teams and the handover to managed operation.\"},\"name\":\"How long does a fixed-scope AI pilot for candidate screening take?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Yes, but only if the data stays within the UK or EEA. OpenAI and Anthropic APIs process data in US or EU regions depending on the tier, which may not satisfy a strict data-residency requirement. For regulated HR data, the safer path is an open-weight model (Llama 3 70B or Mistral 8x22B) running on the client's own hardware or a UK-based GPU cloud. The n8n orchestration layer routes requests to the appropriate model based on data sensitivity, so a single workflow can use a commercial API for general queries and the on-prem model for candidate PII.\"},\"name\":\"Can we use OpenAI or Anthropic APIs for candidate screening under GDPR?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The model drafts a structured assessment: a score from 0 to 100 against each weighted criterion, a short rationale citing the specific line in the CV or Notion page that triggered the score, and a flag for any missing data. A recruiter reviews the draft, can override any score, and must click approve before the candidate moves to the next stage. Every approval or override is logged with a timestamp and the recruiter's ID. The error rate is measured as the percentage of candidates whose final human decision differed from the model's initial draft, tracked weekly during the pilot.\"},\"name\":\"What does the human-in-the-loop approval look like in practice?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The n8n workflow triggers on a new candidate record in the ATS. It fetches the CV, the role brief from Notion or Confluence, and the scoring rubric. The AI model scores the candidate against each criterion and writes the draft assessment back to the ATS. A notification goes to the assigned recruiter. If the recruiter approves within 24 hours, the candidate advances. If they override, the override reason is logged and fed back into the scoring rubric for the next cycle. The entire loop runs in under 90 seconds from trigger to recruiter notification.\"},\"name\":\"How does the n8n orchestration layer handle the candidate screening workflow?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot ships with a measured baseline: the current average time from application to first recruiter decision, and the percentage of candidates who were later found to have been mis-screened (either wrongly rejected or wrongly advanced). After 4 weeks of pilot operation, the same metrics are re-measured. A typical result for a 51-200 person logistics firm is a 40 to 60 percent reduction in screening cycle time and a 25 to 40 percent reduction in mis-screening error rate. These numbers are documented in the pilot report and become the contractual baseline for the rollout phase.\"},\"name\":\"What does the before\/after baseline measurement look like?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The scoring rubric lives in Notion or Confluence as a structured document: each criterion, its weight, the scoring bands, and examples of what a 90 versus a 50 looks like. The n8n workflow reads this document at the start of each screening cycle, so updating the rubric in Notion immediately changes the model's behaviour without any code change. This keeps the HR team in control of the evaluation logic. The model-agnostic architecture means the rubric is model-agnostic too: the same Notion document works whether the backend is GPT-4o, Claude 3.5 Sonnet, or a local Llama 3 instance.\"},\"name\":\"How do we keep the scoring criteria in sync with the AI model?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The model-agnostic design means the n8n workflow calls a model-adapter layer. For general document summarisation, it routes to the OpenAI or Anthropic API. For candidate PII processing, it routes to the open-weight model on the client's hardware. The adapter is a thin HTTP wrapper, so switching models is a configuration change, not a code rewrite. This matters for cost: commercial APIs charge per token, so a high-volume screening operation can cost EUR 200 to 500 per month in API fees. An on-prem Llama 3 70B on a single A100 GPU costs roughly EUR 1,200 per month in cloud GPU rental, but has no per-token fee, making it cheaper above roughly 50,000 tokens per day.\"},\"name\":\"What does model-agnostic mean in practice for cost and switching?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot covers one hiring team and one role family, typically 30 to 80 candidates per month. The rollout extends the same n8n workflow to all hiring teams, adds the round-the-clock response layer for candidate communications, and introduces the predictive scoring model that flags which candidates are most likely to accept an offer based on historical data. The managed operation phase includes monthly model performance reviews, rubric updates, and a 4-hour SLA for any workflow failure. The 6-month timeline assumes the pilot completes by week 12 and the rollout runs weeks 13 to 24.\"},\"name\":\"What does the rollout phase look like after the pilot?\"}]},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/candidate-screening-ai-uk-logistics-n8n-pilot\/#breadcrumbs\",\"@type\":\"BreadcrumbList\",\"itemListElement\":[{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\",\"name\":\"Home\",\"position\":1},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/\",\"name\":\"Blog\",\"position\":2},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/candidate-screening-ai-uk-logistics-n8n-pilot\/\",\"name\":\"Candidate Screening AI for UK Logistics: n8n Pilot vs. Full Rollout\",\"position\":3}]},{\"@id\":\"https:\/\/blog.forfis.com#org\",\"@type\":\"Organization\",\"name\":\"Forfis\",\"url\":\"https:\/\/blog.forfis.com\"}]}","geo_content_hash":"f38e628bb513359a7cf32a0c31be4f6aa738e3a17eb1cd9b874f9fd4229784f6","footnotes":""},"categories":[29],"tags":[71,49,19],"class_list":["post-81","post","type-post","status-publish","format-standard","hentry","category-logistics-and-supply-chain","tag-candidate-screening","tag-reduce-error-rate-in-the-back-office","tag-uk"],"_links":{"self":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/81","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/comments?post=81"}],"version-history":[{"count":0,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/81\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/media?parent=81"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/categories?post=81"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/tags?post=81"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}