8-Week RAG Candidate Screening Pilot for a German E-commerce Team

The problem: manual screening and reporting eat your HR team’s week

You run an e-commerce or retail operation in Germany with 11 to 50 employees. Your HR and recruiting team spends 6 to 10 hours per week manually screening CVs, extracting skills and experience into a spreadsheet, and matching candidates against job postings. The monthly reporting cycle compounds the problem: you pull data from the ATS, reconcile it with the spreadsheet, and format a report for leadership, all by hand. The goal is not to replace the recruiter but to cut the manual back-office work around screening and reporting, so the team spends time on interviews and hiring decisions instead of data entry. The constraint is that candidate data is personal data under GDPR, and your ISO 27001 certification requires documented access controls and audit trails. The pilot must prove a measurable reduction in cycle time and error rate within 8 weeks, using the OpenAI API for the model layer and a custom REST API with webhooks to connect to your existing ATS and reporting tools.

Prerequisites before week one

Before the pilot starts, confirm the following are in place:

  • A working ATS or candidate log. Even a structured spreadsheet with columns for name, email, skills, experience, and job applied to qualifies. The pipeline needs a defined schema to write results back to.
  • A set of 10 to 30 active job postings with written competency requirements. These become the RAG index source. If your job descriptions are vague, the model will match vaguely.
  • A named data owner who can approve the data-processing agreement for the OpenAI API and sign off on the ISO 27001 security annex.
  • A 200-sample gold set of past CVs with manually verified extraction fields. This is your error-rate baseline. Without it, you cannot measure whether the pipeline is accurate.
  • API access to your ATS or reporting tool, or a willingness to expose a minimal REST endpoint. The pilot integrates through custom REST API and webhooks, not by replacing your existing system.
  • A point of contact who can approve scope changes within 48 hours. Fixed-scope means the SOW is locked after week one; slow approvals stall the timeline.

Step 1: Run the process audit and capture the baseline

Spend the first five business days mapping the current workflow. Have the HR team process a sample batch of 50 CVs manually and time each step: receipt, initial read, field extraction, matching against the job posting, and entry into the log. Record the cycle time in minutes per CV and the error rate by having a second person verify the extracted fields. This baseline is the denominator for every metric in the week-8 report. Simultaneously, inventory the document types you receive: PDFs, DOCX, scanned images, and email attachments. Note which fields vary by job type. The audit output is a one-page process map with timestamps and a list of the top five error categories. This document becomes the scope anchor for the pilot SOW.

Step 2: Build the document extraction pipeline

Build the extraction pipeline to parse incoming CVs into structured JSON. Use a document parser such as Apache Tika or a cloud OCR service for scanned PDFs, then feed the text to the OpenAI API with a system prompt that specifies the target schema: name, email, phone, skills (array), years_experience (number), education (array of objects), and job_titles (array). The prompt should include two or three few-shot examples from your gold set to anchor the output format. Log every API call with the input hash, the model version, the response, and a timestamp. Store the structured output in a staging table. The pipeline should handle a batch of 20 CVs in under 90 seconds at the OpenAI gpt-4o token rate, which is roughly 120 tokens per CV for a typical one-page document. If a CV fails to parse, flag it for manual review rather than guessing.

Step 3: Build the RAG index over your job postings

Index your job postings, competency matrices, and past hiring decisions into a vector store. Use a chunking strategy that keeps each job requirement as a separate chunk so the RAG retrieval can cite specific criteria. Embed the chunks with a model such as text-embedding-3-small from OpenAI and store them in a vector database like Weaviate or Qdrant running on your own infrastructure, since the job-posting data may contain internal compensation bands or hiring criteria you do not want in a third-party vector service. The RAG query flow is: take the extracted candidate profile, generate a query string, retrieve the top 5 most relevant job-requirement chunks, and pass them to the OpenAI API with a prompt that asks the model to score the match from 0 to 100 and cite which specific requirements were met or missed. The output is a JSON object with the score, the cited requirements, and a one-paragraph rationale.

Step 4: Wire the REST API and webhooks to your ATS

Expose three REST endpoints: POST /documents to upload a CV, GET /jobs/{id} to retrieve a job posting’s indexed criteria, and POST /results to submit the classification back to your ATS. Configure webhooks so that when the pipeline finishes processing a batch, it fires a batch.completed event to your integration layer with a payload containing the correlation ID, the list of candidate references, the average confidence score, and a link to the full output. Your ATS or integration layer acknowledges with a 200 response within 5 seconds. If it does not, the pipeline retries with exponential backoff: 10 seconds, 30 seconds, 90 seconds. After three failed retries, the record is flagged in the review queue with a webhook_failed status. The human-in-the-loop step sits here: a recruiter sees the model’s score, the cited requirements, and the raw CV side-by-side, and clicks approve or reject. Every approval or rejection is logged with the recruiter’s user ID and timestamp for the ISO 27001 audit trail.

Step 5: Run the pilot with human-in-the-loop review

Run the pipeline on a live batch of 50 to 100 CVs over two weeks. The recruiter reviews every classification, and you log each correction: which field was wrong, what the model said, and what the correct value was. At the end of the run, compute the error rate against the gold set and compare it to the baseline from step 1. If the error rate is above 5 percent, identify the top three error categories and adjust the extraction prompt or the RAG retrieval parameters. Common fixes: tighten the few-shot examples, add a negative constraint to the prompt (“do not infer skills that are not explicitly stated”), or increase the number of retrieved chunks from 5 to 8. Re-run the batch after each adjustment. The goal is to bring the error rate under 5 percent and the cycle time under 30 seconds per CV before the week-8 report. Document every prompt change and its effect in a change log.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *