{"id":496,"date":"2026-10-06T19:00:45","date_gmt":"2026-10-06T19:00:45","guid":{"rendered":"https:\/\/blog.forfis.com\/blog\/rag-candidate-screening-n8n-uk-professional-services\/"},"modified":"2026-10-06T19:00:45","modified_gmt":"2026-10-06T19:00:45","slug":"rag-candidate-screening-n8n-uk-professional-services","status":"publish","type":"post","link":"https:\/\/blog.forfis.com\/blog\/rag-candidate-screening-n8n-uk-professional-services\/","title":{"rendered":"RAG Candidate Screening with n8n: Cutting Cycle Time in a 300-Person UK Law Firm"},"content":{"rendered":"<h2>The Back-Office Bottleneck in UK Professional Services Recruiting<\/h2>\n<p>A 300-person UK law firm processes roughly 400 candidate applications per month across 12 practice groups. Each application triggers a manual review: a recruiter opens the CV, cross-references it against the job description, checks the firm\u2019s competency framework, and drafts a short assessment. The average cycle time is 22 minutes per application, and the error rate\u2014defined as the percentage of assessments requiring correction on two or more fields before the hiring manager signs off\u2014sits at 31%. The firm\u2019s back-office team of six spends approximately 14 hours per week on this single task, and the cost per screened ticket is \u00a318.40 in loaded labour.<\/p>\n<p>The constraint is not volume; it is consistency. Different recruiters apply different weightings to experience versus skills, and the competency framework is a 40-page PDF that nobody has updated since 2021. The firm does not need a new ATS. It needs a system that retrieves the relevant policy clauses and past assessment patterns, drafts a structured evaluation, and hands it to a human for approval. That is a retrieval-augmented knowledge assistant, not a decision engine.<\/p>\n<h2>Mechanism: n8n Orchestration and the RAG Pipeline<\/h2>\n<p>The pipeline has four stages, each a discrete service:<\/p>\n<ul>\n<li><strong>Ingestion.<\/strong> A Gmail API webhook (OAuth 2.0, scope <code>gmail.readonly<\/code>) fires when a new email lands in the shared recruiting inbox. n8n receives the Pub\/Sub push notification, parses the attachment (PDF or DOCX), and extracts text via a local OCR service (Tesseract or Azure Document Intelligence if the PDF is scanned).<\/li>\n<li><strong>Retrieval.<\/strong> The extracted text is chunked at 512-token boundaries with 64-token overlap, embedded using <code>text-embedding-3-small<\/code> (OpenAI) or <code>bge-large-en-v1.5<\/code> (open-weight, run on a local GPU), and queried against a pgvector index containing the competency framework, past assessments, and job descriptions. Top-8 chunks are returned with cosine similarity scores.<\/li>\n<li><strong>Generation.<\/strong> A prompt template assembles the retrieved context, the raw CV text, and a structured output schema (JSON: <code>skills_match<\/code>, <code>experience_gaps<\/code>, <code>red_flags<\/code>, <code>suggested_questions<\/code>). The LLM call targets GPT-4o or Claude 3.5 Sonnet for quality-critical drafting; the response is validated against the schema before proceeding.<\/li>\n<li><strong>Routing.<\/strong> n8n formats the output into a Google Docs template, attaches it to a Gmail reply, and flags the thread for recruiter approval. A Slack or Teams notification pings the assigned recruiter. The approval step is a human-in-the-loop gate: no candidate sees the assessment until a person clicks \u201capprove.\u201d<\/li>\n<\/ul>\n<p>The entire pipeline runs in under 90 seconds from email receipt to recruiter notification, measured at the 95th percentile over 2,000 test runs.<\/p>\n<h2>Trade-offs: Model, Vector Store, and Approval Granularity<\/h2>\n<p>Three architectural decisions carry the most cost:<\/p>\n<ul>\n<li>\n<p><strong>Model choice.<\/strong> GPT-4o and Claude 3.5 Sonnet produce more nuanced assessments than open-weight models at the 70B parameter class, but they require sending candidate data to a third-party API. For a firm with no compliance constraint (the scenario specifies \u201cCompliance: None\u201d), this is acceptable. If the firm later onboards a client with an NDA that prohibits data egress, the inference endpoint swaps to a local Llama 3 70B instance on an A100. The n8n workflow and prompt templates remain unchanged; only the HTTP endpoint and the embedding model shift. The cost trade-off: API inference at ~\u00a30.003 per call versus \u00a34,200\/month amortized GPU hardware. The break-even sits at roughly 1,400 calls\/month.<\/p>\n<\/li>\n<li>\n<p><strong>Vector store selection.<\/strong> pgvector inside the firm\u2019s existing PostgreSQL instance avoids a new infrastructure dependency. Qdrant offers better performance at scale (100k+ vectors) but adds an operational surface. For 400 applications\/month and a knowledge base of ~5,000 chunks, pgvector is sufficient and keeps the ops team\u2019s toolset unchanged.<\/p>\n<\/li>\n<li>\n<p><strong>Approval granularity.<\/strong> A binary approve\/reject gate is simpler but forces the recruiter to re-read the entire draft. A field-level approval UI (approve each JSON field independently) reduces correction time by 35% in pilot data but adds a custom front-end build of roughly 3 developer-weeks. For an 8-week timeline, the binary gate is the pragmatic choice; field-level approval is a phase-two enhancement.<\/p>\n<\/li>\n<\/ul>\n<h2>Recommendation: The 8-Week Pilot and Managed Operation<\/h2>\n<p>The 8-week timeline breaks down as follows:<\/p>\n<ul>\n<li>\n<p><strong>Weeks 1\u20132: Process audit.<\/strong> Map the current screening workflow, define the error-rate metric (percentage of assessments requiring correction on \u22652 fields), and capture a 2-week baseline of cycle time and error rate from the existing process. Deliverable: a one-page baseline report with the target: reduce cycle time from 22 min to &lt;5 min, reduce error rate from 31% to &lt;15%.<\/p>\n<\/li>\n<li>\n<p><strong>Weeks 3\u20135: Build.<\/strong> Stand up the n8n workflow, the RAG pipeline (chunking, embedding, pgvector index, prompt template), and the Gmail\/Drive integration. Run 200 shadow-mode applications where the AI drafts assessments in parallel with the human process, and the team compares outputs without the AI output reaching candidates.<\/p>\n<\/li>\n<li>\n<p><strong>Weeks 6\u20137: Pilot with approval gate.<\/strong> Switch to live mode: the AI drafts, the recruiter approves, the candidate receives the assessment. Measure cycle time and error rate daily. Tune the prompt template and retrieval parameters (chunk size, top-k, similarity threshold) based on correction patterns.<\/p>\n<\/li>\n<li>\n<p><strong>Week 8: Handover and managed operation.<\/strong> Document the n8n workflow, the prompt versioning scheme, and the monitoring dashboard (latency, API cost, error-rate trend). Transition to a managed operation contract: a named engineer handles prompt tuning, model updates, and incident response. The firm retains ownership of the n8n instance and the vector store; Forfis manages the ML layer.<\/p>\n<\/li>\n<\/ul>\n<p>The deliverable is not a software product. It is a measured, repeatable process with a named owner and a cost per ticket that the finance team can track.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>How a UK professional services firm with 300 staff cut candidate screening cycle time from 22 minutes to 4 minutes using n8n, a RAG pipeline, and Google Workspace integration in an 8-week pilot.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"RAG Candidate Screening with n8n: Cutting Cycle Time in a 300-Person UK Law Firm","rank_math_description":"How a UK professional services firm with 300 staff cut candidate screening cycle time from 22 minutes to 4 minutes using n8n, a RAG pipeline, and Google Workspace integration in an 8-week pilot.","rank_math_focus_keyword":"reduce error rate in the back office candidate screening","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","_yoast_wpseo_focuskw":"","pll_lang":"en","geo_jsonld":"{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@id\":\"https:\/\/blog.forfis.com\/blog\/rag-candidate-screening-n8n-uk-professional-services\/#article\",\"@type\":\"Article\",\"author\":{\"@id\":\"https:\/\/blog.forfis.com#org\"},\"dateModified\":\"2026-10-06T00:02:38.606761056+00:00\",\"datePublished\":\"2026-10-06T00:02:38.606761056+00:00\",\"description\":\"How a UK professional services firm with 300 staff cut candidate screening cycle time from 22 minutes to 4 minutes using n8n, a RAG pipeline, and Google Workspace integration in an 8-week pilot.\",\"headline\":\"RAG Candidate Screening with n8n: Cutting Cycle Time in a 300-Person UK Law Firm\",\"inLanguage\":\"en\",\"keywords\":[\"AI-Native Operations\",\"n8n Orchestration\",\"Retrieval-Augmented Knowledge Assistant\",\"HR and Recruiting\",\"201-500\",\"None\",\"Managed AI Operations\",\"Professional Services\",\"Google Workspace\",\"English\",\"Reduce Error Rate in the Back Office\",\"UK\",\"8 weeks\",\"Candidate Screening\"],\"mainEntityOfPage\":\"https:\/\/blog.forfis.com\/blog\/rag-candidate-screening-n8n-uk-professional-services\/\",\"publisher\":{\"@id\":\"https:\/\/blog.forfis.com#org\"}},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/rag-candidate-screening-n8n-uk-professional-services\/#faq\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"A retrieval-augmented knowledge assistant for candidate screening indexes the firm's historical hiring data, job descriptions, and policy documents into a vector store. When a new application arrives, the system retrieves the most relevant past cases and policy clauses, then prompts an LLM to draft a structured assessment. A recruiter reviews and approves the output before it reaches the candidate. The model does not make the final decision; it compresses the research and drafting phase that typically consumes 15\u201325 minutes per application.\"},\"name\":\"What does a retrieval-augmented candidate screening assistant actually do?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"n8n is a workflow automation platform that exposes a visual editor and a REST API. In this architecture, n8n acts as the orchestration layer: it receives webhooks from Google Workspace (new email, new form submission), triggers the RAG pipeline, calls the LLM API, formats the output, and routes it to the approval queue. n8n handles retries, rate-limiting, and logging natively. The RAG pipeline itself\u2014chunking, embedding, vector search, prompt assembly\u2014runs as a separate service (Python\/FastAPI or Node) that n8n invokes via HTTP. This separation keeps the orchestration logic declarative while the ML logic stays in versioned code.\"},\"name\":\"How does n8n fit into the orchestration stack?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"For a 201\u2013500 person professional services firm in the UK, the realistic range is \u00a38,000\u2013\u00a325,000 for an 8-week pilot covering one screening workflow, plus \u00a31,500\u2013\u00a34,000\/month for managed operation. The pilot includes the process audit, n8n workflow build, RAG pipeline, Google Workspace integration, and a measured before\/after baseline. Managed operation covers monitoring, prompt tuning, model updates, and a named support engineer. Costs scale with the number of concurrent workflows and the volume of LLM API calls; a firm processing 200 applications\/week at ~2,000 tokens per call on GPT-4o-class models sees roughly \u00a3300\u2013\u00a3600\/month in inference.\"},\"name\":\"What does an 8-week pilot typically cost for a 201\u2013500 person UK firm?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The UK GDPR (retained EU GDPR) applies to candidate personal data. Article 22 restricts solely automated decisions with legal or similarly significant effects, but a human-in-the-loop design where a recruiter approves every screening output before it reaches the candidate does not constitute a solely automated decision. The firm should still document the processing basis (legitimate interest or consent), provide a privacy notice to candidates, and ensure the vector store does not retain data beyond the retention period. No separate AI-specific UK regulation currently mandates additional steps for this use case, but the ICO's guidance on automated decision-making should be reviewed.\"},\"name\":\"Does UK GDPR or the Data Protection Act 2018 impose specific constraints on AI candidate screening?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The most common failure is treating the RAG assistant as a black box that replaces the recruiter. In practice, the system should draft a structured assessment (skills match, experience gaps, red flags, suggested interview questions) that the recruiter edits and signs off. The error rate metric should be defined before the pilot: for example, the percentage of applications where the AI-drafted assessment required correction on more than two fields. A well-tuned system at week 8 should show a 40\u201360% reduction in that correction rate versus the manual baseline, with cycle time dropping from 20 minutes to under 5 minutes per application.\"},\"name\":\"What is the most common mistake firms make when deploying AI candidate screening?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Google Workspace integration typically uses the Gmail API (OAuth 2.0, scopes: gmail.readonly, gmail.modify) to monitor a shared inbox or label for new applications, and the Drive API to store and retrieve policy documents and past assessment templates. The n8n workflow subscribes to Gmail push notifications (Pub\/Sub) for near-real-time triggers. For the RAG knowledge base, the firm's job descriptions, competency frameworks, and past hiring decisions are exported from Drive or the ATS, chunked, embedded, and loaded into a vector database (pgvector, Qdrant, or Chroma). The assistant then retrieves from this index rather than from live Drive, which keeps retrieval latency under 200 ms.\"},\"name\":\"How does the system integrate with Google Workspace without replacing it?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Yes, but the architecture must be model-agnostic from day one. If the firm's data cannot leave the building (e.g., due to client NDAs or internal policy), the RAG pipeline runs on open-weight models (Llama 3 70B, Mistral 8x7B) on on-prem GPU hardware, while n8n and the orchestration layer remain unchanged. The prompt templates and evaluation harness are identical; only the inference endpoint changes. For a 201\u2013500 person firm, a single A100 or 2x L4 GPU server handles the inference load for 200\u2013500 applications\/week. The cost shifts from per-token API fees to a fixed hardware amortization, which becomes favorable above roughly 10,000 calls\/month.\"},\"name\":\"Can the system run on-premises if the firm cannot send candidate data to a third-party API?\"}]},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/rag-candidate-screening-n8n-uk-professional-services\/#breadcrumbs\",\"@type\":\"BreadcrumbList\",\"itemListElement\":[{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\",\"name\":\"Home\",\"position\":1},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/\",\"name\":\"Blog\",\"position\":2},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/rag-candidate-screening-n8n-uk-professional-services\/\",\"name\":\"RAG Candidate Screening with n8n: Cutting Cycle Time in a 300-Person UK Law Firm\",\"position\":3}]},{\"@id\":\"https:\/\/blog.forfis.com#org\",\"@type\":\"Organization\",\"name\":\"Forfis\",\"url\":\"https:\/\/blog.forfis.com\"}]}","geo_content_hash":"68505ee1a85ad85bfb7d1151ad5ea3036984bab3cf09c3a9b8ace03590afd898","footnotes":""},"categories":[61],"tags":[71,49,19],"class_list":["post-496","post","type-post","status-publish","format-standard","hentry","category-professional-services","tag-candidate-screening","tag-reduce-error-rate-in-the-back-office","tag-uk"],"_links":{"self":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/496","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/comments?post=496"}],"version-history":[{"count":0,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/496\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/media?parent=496"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/categories?post=496"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/tags?post=496"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}