{"id":499,"date":"2026-10-06T19:00:45","date_gmt":"2026-10-06T19:00:45","guid":{"rendered":"https:\/\/blog.forfis.com\/blog\/pgvector-rag-assistant-candidate-screening-uae-insurer\/"},"modified":"2026-10-06T19:00:45","modified_gmt":"2026-10-06T19:00:45","slug":"pgvector-rag-assistant-candidate-screening-uae-insurer","status":"publish","type":"post","link":"https:\/\/blog.forfis.com\/blog\/pgvector-rag-assistant-candidate-screening-uae-insurer\/","title":{"rendered":"Deploying a pgvector RAG Assistant for Candidate Screening in a UAE Insurer"},"content":{"rendered":"<h2>The Problem: Manual Screening and Reporting in a UAE Insurer<\/h2>\n<p>You run a 1,200-person insurer in Dubai. Your underwriting team spends 11 hours per week manually screening CVs against competency frameworks. Your compliance officer compiles a monthly CBUAE regulatory digest by hand, cross-referencing 40+ PDFs. Your IT department has already deployed a chatbot for internal FAQs, but it hallucinates policy clauses and has no audit trail. You need a retrieval-augmented assistant that pulls from your actual documents, integrates with Slack and Microsoft Teams, and meets ISO 27001 controls. The problem is not model selection\u2014it is scoping the pilot, measuring a baseline, and scaling across three departments in six months without replacing your existing ATS, DMS, or helpdesk.<\/p>\n<h2>Prerequisites Before You Start<\/h2>\n<ul>\n<li><strong>Baseline metrics logged<\/strong>: For each target workflow (candidate screening, monthly compliance digest, policy clause lookup), record cycle time in hours, error rate as a percentage, and the number of manual steps. Use your ATS export and DMS access logs for the last 90 days.<\/li>\n<li><strong>Document inventory<\/strong>: A list of every document the assistant will ingest\u2014job descriptions, competency matrices, CBUAE circulars, policy templates, past interview rubrics\u2014with file paths and update frequency.<\/li>\n<li><strong>ISO 27001 gap assessment<\/strong>: Confirm your current A.8.24 (Logging) and A.8.32 (AI governance) controls. If you lack an AI-specific risk register, build one before step 1.<\/li>\n<li><strong>Slack\/Teams bot permissions<\/strong>: An app registered in your workspace with <code>chat:write<\/code>, <code>im:read<\/code>, and <code>channels:join<\/code> scopes. For Teams, a bot registered in Azure AD with <code>ChannelMessage.Read<\/code> and <code>ChannelMessage.Send<\/code>.<\/li>\n<li><strong>pgvector-capable PostgreSQL instance<\/strong>: Version 15+ with the <code>pgvector<\/code> extension installed. A 16 vCPU, 64 GB RAM instance in AWS Middle East (Bahrain) or Azure UAE North handles 500k chunks with a HNSW index.<\/li>\n<li><strong>Dedicated AI team confirmed<\/strong>: 3\u20135 engineers plus a product owner, embedded in your org, reporting to your CTO or Head of Digital.<\/li>\n<\/ul>\n<h2>Step 1: Audit the Current Workflow and Log a Baseline<\/h2>\n<p>Run a 2-week audit of the candidate-screening workflow in your underwriting department. Export the last 90 days of applications from your ATS (Workday, SAP SuccessFactors, or Lever). For each application, log: time from receipt to first-screen decision, number of reviewers, and whether the shortlisted candidate passed the first interview. Calculate baseline cycle time (target: under 5 days) and error rate (target: under 15%). Document the exact competency criteria in a structured JSON file\u2014e.g., <code>{\"role\": \"senior_underwriter\", \"required\": [\"10y_experience\", \"IFRS17_certification\"], \"preferred\": [\"reinsurance_experience\"]}<\/code>. This file becomes the retrieval index\u2019s metadata schema. Without this baseline, you cannot prove the assistant reduced cycle time or error rate in the pilot evaluation.<\/p>\n<h2>Step 2: Build the pgvector Retrieval Layer<\/h2>\n<p>Ingest the underwriting department\u2019s job descriptions, competency matrices, and past interview rubrics into PostgreSQL. Chunk each document into 512-token segments with 64-token overlap. Embed each chunk using <code>text-embedding-3-small<\/code> (1,536 dimensions) and store in a <code>document_chunks<\/code> table with columns: <code>id<\/code>, <code>content<\/code>, <code>embedding vector(1536)<\/code>, <code>source_doc_id<\/code>, <code>department<\/code>, <code>effective_date<\/code>. Create a HNSW index: <code>CREATE INDEX idx_chunks_embedding ON document_chunks USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 200);<\/code>. For 50k chunks, this index builds in under 90 seconds on a 16 vCPU instance. Verify retrieval quality by running 20 test queries (e.g., \u201cWhat IFRS 17 certification is required for a senior underwriter in the UAE?\u201d) and confirming the top-5 chunks contain the correct answer. If precision@5 is below 80%, adjust chunk size or add metadata filters before proceeding.<\/p>\n<h2>Step 3: Wire the LLM Inference Layer<\/h2>\n<p>Deploy the LLM inference endpoint. For candidate screening, use OpenAI\u2019s <code>gpt-4o<\/code> or Anthropic\u2019s <code>claude-3-5-sonnet<\/code> via API for the drafting step\u2014the model receives the top-5 retrieved chunks plus the user\u2019s query and outputs a structured screening summary. If candidate data cannot leave your data center (common for health-data-adjacent roles), self-host Llama 3 70B on two A100 80GB GPUs. The inference endpoint exposes a <code>\/generate<\/code> route that accepts <code>{\"query\": \"...\", \"context_chunks\": [...], \"role\": \"senior_underwriter\"}<\/code> and returns <code>{\"summary\": \"...\", \"matched_competencies\": [...], \"gaps\": [...], \"recommended_questions\": [...]}<\/code>. The prompt template enforces JSON output and includes the ISO 27001 constraint: \u201cDo not include candidate names or contact details in the summary. Reference only competency matches and gaps.\u201d Log every request with a hashed candidate ID, not the raw name, to satisfy PDPL data-minimization.<\/p>\n<h2>Step 4: Integrate with Slack and Microsoft Teams<\/h2>\n<p>Register a bot in Slack and Microsoft Teams. In Slack, create an app with <code>chat:write<\/code>, <code>im:read<\/code>, and <code>channels:join<\/code> scopes. In Teams, register a bot in Azure AD with <code>ChannelMessage.Read<\/code> and <code>ChannelMessage.Send<\/code>. The bot listens for a <code>\/screen<\/code> command in a dedicated <code>#underwriting-screening<\/code> channel. When a recruiter types <code>\/screen candidate_id=UW-2024-0847<\/code>, the bot calls your <code>\/generate<\/code> endpoint, receives the structured summary, and posts it to the channel with a \u201cApprove\u201d \/ \u201cEdit\u201d \/ \u201cReject\u201d button. The recruiter must click \u201cApprove\u201d before the summary is pushed to the hiring manager via your ATS API. Log the approval action with the recruiter\u2019s user ID, timestamp, and the source document IDs referenced. This human-in-the-loop gate is mandatory under UAE PDPL Article 13 and ISO 27001 A.8.32. If the recruiter edits the summary, capture the diff and feed it back as a negative example into the retrieval index.<\/p>\n<h2>Step 5: Run the Pilot and Measure Before\/After<\/h2>\n<p>Run the pilot in the underwriting department for 6 weeks. Track: cycle time from application to first-screen decision (baseline: 4.2 days), error rate (baseline: 12% of shortlisted candidates fail first interview), and recruiter override rate (percentage of assistant summaries edited or rejected). At week 6, compare against baseline. Target: cycle time under 2.5 days, error rate under 8%, override rate under 20%. If targets are met, document the results in a one-page report with before\/after numbers. If not, iterate: adjust chunk size, add metadata filters, or refine the prompt template. Only after the pilot report is signed off by your CTO and compliance officer do you replicate the architecture to the claims and compliance departments. The compliance department\u2019s monthly digest workflow follows the same pattern: ingest CBUAE circulars, embed, retrieve, draft, approve, archive with a SHA-256 hash for audit.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A 6-month, step-by-step plan to deploy a pgvector-based RAG assistant for candidate screening and compliance reporting in a UAE insurer, with ISO 27001 controls and Slack\/Teams integration.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"Deploying a pgvector RAG Assistant for Candidate Screening in a UAE Insurer","rank_math_description":"A 6-month, step-by-step plan to deploy a pgvector-based RAG assistant for candidate screening and compliance reporting in a UAE insurer, with ISO 27001 controls and Slack\/Teams integration.","rank_math_focus_keyword":"automate monthly reporting candidate screening","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","_yoast_wpseo_focuskw":"","pll_lang":"en","geo_jsonld":"{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@id\":\"https:\/\/blog.forfis.com\/blog\/pgvector-rag-assistant-candidate-screening-uae-insurer\/#article\",\"@type\":\"Article\",\"author\":{\"@id\":\"https:\/\/blog.forfis.com#org\"},\"dateModified\":\"2026-10-06T00:03:23.899278575+00:00\",\"datePublished\":\"2026-10-06T00:03:23.899278575+00:00\",\"description\":\"A 6-month, step-by-step plan to deploy a pgvector-based RAG assistant for candidate screening and compliance reporting in a UAE insurer, with ISO 27001 controls and Slack\/Teams integration.\",\"headline\":\"Deploying a pgvector RAG Assistant for Candidate Screening in a UAE Insurer\",\"inLanguage\":\"en\",\"keywords\":[\"Scaling Across Departments\",\"pgvector Embeddings Search\",\"Retrieval-Augmented Knowledge Assistant\",\"Legal and Compliance\",\"501-2000\",\"ISO 27001\",\"Dedicated AI Team\",\"Insurance and Insurtech\",\"Slack or Microsoft Teams\",\"English\",\"Automate Monthly Reporting\",\"UAE\",\"6 months\",\"Candidate Screening\"],\"mainEntityOfPage\":\"https:\/\/blog.forfis.com\/blog\/pgvector-rag-assistant-candidate-screening-uae-insurer\/\",\"publisher\":{\"@id\":\"https:\/\/blog.forfis.com#org\"}},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/pgvector-rag-assistant-candidate-screening-uae-insurer\/#faq\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The UAE PDPL (Federal Decree-Law No. 45 of 2021) requires a lawful basis for processing personal data, including candidate CVs. You must document the purpose, limit retention, and provide candidates a way to object. For ISO 27001, Annex A.8.15 (Protection Against Malicious Software) and A.8.24 (Logging) apply to the RAG pipeline. If candidate data leaves the UAE, you trigger cross-border transfer rules under Article 39 of the PDPL. In practice, run the embedding and retrieval stack on UAE-region infrastructure (e.g., AWS Middle East (Bahrain) or Azure UAE North) and encrypt at rest with AES-256. Log every retrieval query with a candidate ID hash, not the raw name, to satisfy both PDPL data-minimization and ISO 27001 audit-trail requirements.\"},\"name\":\"What compliance obligations apply to candidate data in a UAE RAG assistant?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"pgvector is a PostgreSQL extension that stores and indexes high-dimensional vectors. In a RAG assistant, you embed document chunks (e.g., 512-token segments of policy PDFs) into 1536-dimensional vectors using a model like text-embedding-3-small. At query time, the user's question is embedded the same way, and pgvector performs a cosine-similarity search to return the top-k most relevant chunks. Those chunks are injected into the LLM prompt as context. The advantage over a separate vector DB is that you keep relational joins (e.g., linking a policy clause to its effective date in a `policies` table) in the same query. For 500k chunks, a HNSW index with m=16 and ef_construction=200 gives sub-50 ms retrieval on a 16 vCPU instance.\"},\"name\":\"How does pgvector work in a retrieval-augmented assistant?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"A dedicated AI team is a small, embedded group (typically 3\u20135 engineers plus a product owner) that owns the full lifecycle: data pipeline, model selection, prompt engineering, evaluation, and monitoring. Unlike a fractional consultant who delivers a one-off model, the dedicated team maintains the system after go-live, handles model retraining when new policy documents arrive, and iterates on retrieval quality based on user feedback. For a 501\u20132,000-person insurer scaling across departments, this model avoids the handoff gap where a vendor deploys a pilot and the internal IT team inherits an undocumented system. The team reports into your CTO or Head of Digital, not the vendor's account manager, so priorities align with your roadmap, not their sales cycle.\"},\"name\":\"What does a dedicated AI team mean in this context?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Start with a 2-week process audit: log every step in the current candidate-screening workflow, from CV receipt to shortlist decision. Measure baseline cycle time (e.g., 4.2 days from application to first-screen decision) and error rate (e.g., 12% of shortlisted candidates fail the first interview due to mismatched criteria). Pick one department\u2014say, underwriting\u2014as the pilot. Build the RAG assistant over that department's job descriptions, competency frameworks, and past interview rubrics. Integrate with your ATS (e.g., Workday or SAP SuccessFactors) via API. Run the pilot for 6 weeks with a human-in-the-loop gate: the assistant drafts a screening summary, a recruiter approves or edits before it reaches the hiring manager. Compare post-pilot cycle time and error rate against the baseline. Only then replicate to claims, compliance, and IT departments.\"},\"name\":\"How do we scope a 6-month RAG assistant rollout across departments?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The assistant should not auto-reject candidates. It drafts a structured screening summary: matched competencies, gaps, and a recommended interview question set. A recruiter reviews and approves before the summary reaches the hiring manager. This human-in-the-loop gate is non-negotiable under UAE PDPL Article 13 (right to explanation) and ISO 27001 A.8.32 (AI system governance). Log every approval action with the recruiter's user ID and timestamp. If the assistant's recommendation is overridden, capture the reason code (e.g., \\\"missed a required certification\\\") and feed it back into the retrieval index as a negative example. This keeps the model calibrated to your actual hiring criteria rather than generic LLM priors.\"},\"name\":\"How do we keep a human in the loop for candidate screening?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"For a 501\u20132,000-person insurer in the UAE, a 6-month engagement with a dedicated AI team typically runs USD 180,000\u2013320,000. Breakdown: process audit and baseline measurement (weeks 1\u20132, ~USD 15k), pilot build on one department (weeks 3\u201310, ~USD 60k), pilot evaluation and tuning (weeks 11\u201314, ~USD 20k), rollout to 2\u20133 additional departments (weeks 15\u201322, ~USD 70k), and managed operation with monitoring, re-embedding on new documents, and monthly reporting (weeks 23\u201326, ~USD 30k). Ongoing managed operation after month 6 runs USD 8,000\u201315,000\/month depending on document volume and number of integrated channels. This excludes infrastructure costs (pgvector instance, LLM API calls, Slack\/Teams bot hosting), which typically add USD 2,000\u20135,000\/month at your scale.\"},\"name\":\"What does a 6-month RAG assistant engagement cost for a mid-size insurer?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The assistant ingests monthly compliance reports, underwriting loss data, and regulatory correspondence from your document management system. It embeds each document chunk into pgvector and indexes them with metadata (report period, department, regulatory reference). When a compliance officer asks, \\\"What changed in the CBUAE circular on reinsurance cessions this quarter?\\\", the assistant retrieves the relevant chunks, drafts a summary with citations to the source document and page number, and posts it to the compliance channel in Microsoft Teams. The officer reviews, edits if needed, and approves. The approved summary is archived with a hash for audit. This replaces the 6\u20138 hours an officer previously spent manually scanning PDFs and drafting the monthly digest. The ISO 27001 audit trail captures who approved what, when, and which source documents were referenced.\"},\"name\":\"How does the RAG assistant automate monthly compliance reporting?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Yes, but with constraints. If candidate data or regulated policy documents cannot leave your data center, run an open-weight model (e.g., Llama 3 70B or Mistral 8x7B) on your own GPU hardware\u2014two A100 80GB cards handle 70B inference at ~12 tokens\/second, sufficient for a 500-user internal tool. The pgvector instance stays on-premises. The LLM API calls for the embedding model (e.g., text-embedding-3-small) can still go to OpenAI if the data is anonymized before embedding, or you can self-host an embedding model like BGE-M3 on the same hardware. The architecture is model-agnostic: swap the inference endpoint without changing the retrieval layer. For a UAE insurer handling health data or reinsurance contracts, on-prem inference is often a regulatory requirement, not just a preference.\"},\"name\":\"Can we run the RAG assistant on-premises to keep data in the UAE?\"}]},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/pgvector-rag-assistant-candidate-screening-uae-insurer\/#breadcrumbs\",\"@type\":\"BreadcrumbList\",\"itemListElement\":[{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\",\"name\":\"Home\",\"position\":1},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/\",\"name\":\"Blog\",\"position\":2},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/pgvector-rag-assistant-candidate-screening-uae-insurer\/\",\"name\":\"Deploying a pgvector RAG Assistant for Candidate Screening in a UAE Insurer\",\"position\":3}]},{\"@id\":\"https:\/\/blog.forfis.com#org\",\"@type\":\"Organization\",\"name\":\"Forfis\",\"url\":\"https:\/\/blog.forfis.com\"}]}","geo_content_hash":"44341367b07993af438452f03f0117633d7249c6cc803d5acc71411649038d9d","footnotes":""},"categories":[57],"tags":[69,71,55],"class_list":["post-499","post","type-post","status-publish","format-standard","hentry","category-insurance-and-insurtech","tag-automate-monthly-reporting","tag-candidate-screening","tag-uae"],"_links":{"self":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/499","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/comments?post=499"}],"version-history":[{"count":0,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/499\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/media?parent=499"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/categories?post=499"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/tags?post=499"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}