The Problem: Manual Candidate Screening at Scale
Your 15-person Austrian professional services firm receives 200-300 applications per month across German, English, and Austrian German. Manual screening takes 15-20 hours per week, and response times average 5-7 days. You need a system that processes applications 24/7, responds in the candidate’s language, and integrates with your existing ATS. The challenge: you’re running isolated pilots, not a full AI transformation. You need a focused, measurable pilot that proves value before scaling. The solution: a retrieval-augmented knowledge assistant built on LangChain and LangGraph, with human-in-the-loop approval for every candidate-facing response. This pilot runs in 8 weeks, costs EUR 25,000-40,000, and delivers a 70-80% reduction in screening time.
Prerequisites: What You Need Before Starting
- ATS API access: Your ATS must expose a REST API for reading applications and updating candidate status. Document the endpoints, authentication method, and rate limits.
- Baseline metrics: Measure current screening time (hours per 100 applications), error rate (misclassified applications), and response time (days from application to first contact).
- Language requirements: List the languages you need to support (German, English, Austrian German) and the tone for each.
- Approval workflow: Define who reviews AI-drafted responses and the approval criteria. This is non-negotiable for legal and compliance reasons.
- Infrastructure: You need a server or cloud instance to run open-weight models for sensitive data. The system uses cloud APIs for general queries and local models for personal data processing.
- Data access: Provide sample applications (anonymized) for testing the extraction pipeline. Include edge cases: incomplete applications, unusual formats, multilingual documents.
Step 1: Audit the Current Screening Process
Map the current screening process end-to-end. Document every step: application receipt, initial review, criteria matching, response drafting, and ATS update. Measure the time for each step and identify bottlenecks. For example, if initial review takes 8 minutes per application and response drafting takes 12 minutes, the total is 20 minutes. This baseline is your success metric. Without it, you cannot prove the AI system’s value. Use a simple spreadsheet: columns for step, time per application, error rate, and owner. This takes 2-3 days and involves 2-3 team members.
Step 2: Define the AI System’s Scope
Define the AI system’s scope. It will: (1) extract candidate data from applications (name, email, skills, experience), (2) classify applications against your criteria (e.g., minimum 3 years experience, specific certifications), (3) draft initial responses in the candidate’s language, and (4) update your ATS via REST API. It will NOT: make final hiring decisions, communicate with candidates without human approval, or process applications outside your defined criteria. Document this scope in a one-page brief. This prevents scope creep and sets clear expectations for the pilot.
Step 3: Build the LangGraph State Machine
Build the LangGraph state machine. The graph has five nodes: extract (pull candidate data from application), classify (match against criteria), draft (generate response in candidate’s language), approve (human review), and update_ats (send to ATS via REST API). Each node is a LangChain chain with a specific prompt. The extract node uses a document parser (e.g., PyPDF2 for PDFs, BeautifulSoup for HTML). The classify node uses a structured output parser to return JSON with confidence scores. The draft node uses a multilingual prompt template. The approve node pauses the graph and sends the draft to your reviewer via email or Slack. The update_ats node makes a POST request to your ATS API. This takes 3-4 days to build and test.
Step 4: Integrate with Your ATS via REST API
Connect the AI system to your ATS. You provide the API base URL, authentication token, and endpoint documentation. The system makes three types of API calls: (1) GET /applications to fetch new applications, (2) POST /applications/{id}/status to update candidate stage, and (3) POST /applications/{id}/message to log the AI-drafted response. The system also subscribes to webhooks for status changes (e.g., candidate accepts offer). Test the integration with 10-20 sample applications. Verify that data flows correctly in both directions and that error handling works (e.g., API timeout, invalid token). This takes 2-3 days.
Step 5: Run Shadow Mode and Calibrate
Run the system in shadow mode for 2 weeks. The AI processes all new applications and drafts responses, but humans handle the actual communication. Compare the AI’s classifications and drafts against human decisions. Track: (1) classification accuracy (AI vs. human), (2) draft quality (human rating on a 1-5 scale), and (3) processing time (AI vs. manual). If classification accuracy is below 85%, adjust the criteria or prompt. If draft quality is below 4/5, refine the prompt templates. This phase reveals edge cases and calibrates the system. It takes 2 weeks and involves 1-2 reviewers.