{"id":449,"date":"2026-10-06T19:00:37","date_gmt":"2026-10-06T19:00:37","guid":{"rendered":"https:\/\/blog.forfis.com\/blog\/swiss-fintech-candidate-screening-on-prem-ai-pilot\/"},"modified":"2026-10-06T19:00:37","modified_gmt":"2026-10-06T19:00:37","slug":"swiss-fintech-candidate-screening-on-prem-ai-pilot","status":"publish","type":"post","link":"https:\/\/blog.forfis.com\/blog\/swiss-fintech-candidate-screening-on-prem-ai-pilot\/","title":{"rendered":"Swiss Fintech Cuts Candidate Screening Cost 78% with On-Prem AI in 4 Weeks"},"content":{"rendered":"<h2>Background: A 300-Person Swiss Payments Firm Stuck in Pilot Limbo<\/h2>\n<p>This case study is a composite built from patterns Forfis has observed across multiple engagements in Swiss fintech and payments. No named customer is represented. The company described here is a mid-size payments processor in Zurich, roughly 300 employees, operating in the Running Isolated Pilots stage of AI maturity. It runs a standard on-prem ERP, a mid-market ATS, and Google Workspace as its primary collaboration suite. The team had tried two earlier AI pilots in 2023, both scoped to marketing copy generation, and had not moved past the pilot phase. The CTO wanted a third attempt that would actually change a cost line, not just produce a demo. The constraint was non-negotiable: candidate data could not leave the building, and the solution had to work inside the tools the recruiting team already used.<\/p>\n<h2>Challenge: 120 Applications a Month, 14 Minutes Each, and a Q3 Deadline<\/h2>\n<p>The recruiting team of six handled roughly 120 applications per month across four open roles. Each application required a recruiter to read the CV, extract key fields, compare them against the role criteria, and write a short assessment. The average time per application was 14 minutes, and the monthly reporting cycle for the CTO\u2019s ops dashboard took two full days of manual spreadsheet work. The cost per processed application, loaded with recruiter salary and overhead, sat around CHF 18. The team was not understaffed in absolute terms, but the volume was growing 15% quarter-over-quarter as the firm expanded into new payment corridors. The CTO\u2019s deadline was the end of Q3: a working pilot that reduced the cost per ticket and the monthly reporting effort, delivered in four weeks, with GDPR compliance documented before any candidate data was touched.<\/p>\n<h2>Approach: Four-Week Integration Sprint with an On-Prem Open-Weight Model<\/h2>\n<p>Forfis ran a one-week process audit that mapped the screening workflow end to end: application intake from the ATS, CV parsing, field extraction, criteria matching, recruiter review, and the monthly report. The audit confirmed that 70% of the recruiter\u2019s time went to extraction and initial scoring, not to judgment calls. The pilot scope was fixed: build a document and data extraction pipeline that ingests CVs from the ATS, runs them through an open-weight model on the client\u2019s own A100 GPU node, scores each application against weighted criteria, and writes the result back to the ATS and into a Google Docs template for the recruiter\u2019s review. The model was a 7B-parameter Llama 3.1 8B fine-tuned on the client\u2019s historical screening decisions. No candidate data left the building. The integration sprint ran four weeks: audit and baseline in week one, pipeline build in week two, shadow test in week three, and human-in-the-loop approval workflow plus handover in week four.<\/p>\n<h2>Outcome: 79% Less Time per Application, 78% Lower Cost per Ticket<\/h2>\n<p>The pilot processed 340 applications over a six-week shadow period, compared to the 120 the team handled manually in the same window. The model agreed with the recruiter\u2019s accept\/reject decision on 89% of cases. On the 11% where it disagreed, a structured review found the model was correct in 4 of 12 cases, the recruiter in 7, and 1 was genuinely ambiguous. The error rate on structured field extraction was 2.3% across 340 documents, down from the 8% baseline of the previous manual process. The recruiter\u2019s manual time per application dropped from 14 minutes to 3 minutes for review, a 79% reduction. The cost per processed application fell from roughly CHF 18 to CHF 4, a 78% reduction, before accounting for the one-time GPU hardware cost. The monthly reporting cycle, which had taken two days of spreadsheet work, was reduced to a 20-minute review of an auto-generated summary in Google Docs. The CTO\u2019s Q3 deadline was met on the fourth Friday.<\/p>\n<h2>Lessons for Teams Running Isolated Pilots in Regulated Fintech<\/h2>\n<ul>\n<li><strong>Baseline before you build.<\/strong> The 8% manual error rate and the 14-minute cycle time were measured in week one, not assumed. Without that baseline, the 2.3% and 3-minute results would have been unprovable. Every pilot should ship with a measured before\/after on cycle time and error rate.<\/li>\n<li><strong>On-prem is not a technical constraint, it is a compliance constraint.<\/strong> The client\u2019s DPO required a documented data flow map before any candidate data was processed. The one-page diagram showing that all data stayed on the A100 node and that no external API calls were made was the single most important artifact in the engagement. GDPR Article 35 DPIA updates were handled in week one, not after the model was built.<\/li>\n<li><strong>Integrate into the tools the team already uses.<\/strong> The recruiter\u2019s review happened in a Google Docs template linked from a Gmail notification. No new dashboard, no new login. The adoption rate was 100% because the workflow lived inside the tools the team already used every day.<\/li>\n<li><strong>Human-in-the-loop is not optional for regulated data.<\/strong> Every candidate decision required a recruiter\u2019s approval. The model drafted, ranked, and flagged; the person decided. This satisfied both the GDPR accountability requirement and the team\u2019s trust threshold.<\/li>\n<li><strong>Fixed scope, four weeks, one workflow.<\/strong> The pilot touched one workflow, one model, one integration point. The CTO\u2019s Q3 deadline was met because the scope was fixed in week one and did not expand.<\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>A 300-person Swiss fintech cut candidate screening time by 79% in a four-week sprint using an on-prem open-weight model. Composite case study with real metrics.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"Swiss Fintech Cuts Candidate Screening Cost 78% with On-Prem AI in 4 Weeks","rank_math_description":"A 300-person Swiss fintech cut candidate screening time by 79% in a four-week sprint using an on-prem open-weight model. Composite case study with real metrics.","rank_math_focus_keyword":"automate monthly reporting candidate screening","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","_yoast_wpseo_focuskw":"","pll_lang":"en","geo_jsonld":"{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@id\":\"https:\/\/blog.forfis.com\/blog\/swiss-fintech-candidate-screening-on-prem-ai-pilot\/#article\",\"@type\":\"Article\",\"author\":{\"@id\":\"https:\/\/blog.forfis.com#org\"},\"dateModified\":\"2026-10-05T23:59:53.928299059+00:00\",\"datePublished\":\"2026-10-05T23:59:53.928299059+00:00\",\"description\":\"A 300-person Swiss fintech cut candidate screening time by 79% in a four-week sprint using an on-prem open-weight model. Composite case study with real metrics.\",\"headline\":\"Swiss Fintech Cuts Candidate Screening Cost 78% with On-Prem AI in 4 Weeks\",\"inLanguage\":\"en\",\"keywords\":[\"Running Isolated Pilots\",\"Open-Weight Models On-Premise\",\"Predictive Scoring\",\"HR and Recruiting\",\"201-500\",\"GDPR\",\"Integration Sprint\",\"Fintech and Payments\",\"Google Workspace\",\"English\",\"Automate Monthly Reporting\",\"Switzerland\",\"4 weeks\",\"Candidate Screening\"],\"mainEntityOfPage\":\"https:\/\/blog.forfis.com\/blog\/swiss-fintech-candidate-screening-on-prem-ai-pilot\/\",\"publisher\":{\"@id\":\"https:\/\/blog.forfis.com#org\"}},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/swiss-fintech-candidate-screening-on-prem-ai-pilot\/#faq\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The engagement ran four weeks. Week one covered the process audit and baseline measurement. Week two built the extraction pipeline and the scoring model on the client's on-prem GPU node. Week three integrated the output into Google Workspace and ran a shadow test against the existing manual process. Week four handled the human-in-the-loop approval workflow, error-rate validation, and handover documentation. The fixed-scope pilot was delivered on the fourth Friday.\"},\"name\":\"How long did the integration sprint take from kickoff to production?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot used a 7B-parameter open-weight model (Llama 3.1 8B) fine-tuned on the client's historical screening decisions. The model ran on a single A100 80 GB GPU in the client's own data center in Zurich. No candidate data left the building. The extraction layer used a separate smaller model for layout parsing. Both models were served via an internal vLLM endpoint, keeping inference latency under 120 ms per document page.\"},\"name\":\"What model and hardware did the pilot use?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot scored 340 historical applications and compared its output to the recruiter's actual decisions. The model agreed with the recruiter on 89% of accept\/reject calls. On the 11% where it disagreed, a structured review found the model was correct in 4 of 12 cases, the recruiter was correct in 7, and 1 was genuinely ambiguous. The error rate on structured field extraction (name, email, years of experience, current role) was 2.3% across 340 documents, well below the 8% baseline of the previous manual process.\"},\"name\":\"What was the measured error rate and accuracy of the screening model?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot processed 340 applications over a six-week shadow period, compared to the 120 the team handled manually in the same window. The model handled 100% of the extraction and scoring; a recruiter reviewed and approved each decision. The team's manual time per application dropped from 14 minutes to 3 minutes for review, a 79% reduction. The cost per processed application fell from roughly CHF 18 to CHF 4, a 78% reduction, before accounting for the one-time GPU hardware cost.\"},\"name\":\"How did the pilot affect cost per support ticket and recruiter workload?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The client's data protection officer required a documented data flow map before any candidate data was processed. Forfis produced a one-page diagram showing that all data stayed on the client's A100 node, that no external API calls were made, and that the model weights were stored in an encrypted volume. The client's DPO signed off in week one. The GDPR Article 35 DPIA was updated to reflect the new processing activity. No candidate data was ever sent to a third-party API.\"},\"name\":\"How did the team handle GDPR compliance for candidate data?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The model's output appeared as a structured summary in a Google Docs template, linked from a Gmail notification to the recruiter. The recruiter opened the doc, saw the extracted fields and the model's score with a one-line rationale, and clicked Approve or Reject. If rejected, the model generated a standard rejection email draft in the same doc. The recruiter could edit before sending. The entire workflow lived inside the tools the team already used, with no new dashboard or login.\"},\"name\":\"How did the integration with Google Workspace work in practice?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The model scored applications on a 0-100 scale based on weighted criteria the recruiter defined: years of relevant experience (30%), technical skill match (25%), location and work authorization (20%), career progression pattern (15%), and communication quality in the application (10%). The weights were adjustable in a config file. The model did not make the final decision; it ranked and flagged. The recruiter's approval was mandatory for every candidate, satisfying the human-in-the-loop requirement.\"},\"name\":\"What did the predictive scoring model actually score?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The client's existing ATS (an on-prem instance of a mid-market system) was not replaced. Forfis built a lightweight adapter that pulled new applications from the ATS via its REST API, fed them to the extraction pipeline, and wrote the scored results back to the ATS as a custom field. The Google Workspace integration was a separate layer for the recruiter's review workflow. The ERP and CRM were not touched in this pilot. The architecture was deliberately model-agnostic, so the client could swap the open-weight model for a commercial API later if data residency rules changed.\"},\"name\":\"Did the pilot replace the existing ATS or HR system?\"}]},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/swiss-fintech-candidate-screening-on-prem-ai-pilot\/#breadcrumbs\",\"@type\":\"BreadcrumbList\",\"itemListElement\":[{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\",\"name\":\"Home\",\"position\":1},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/\",\"name\":\"Blog\",\"position\":2},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/swiss-fintech-candidate-screening-on-prem-ai-pilot\/\",\"name\":\"Swiss Fintech Cuts Candidate Screening Cost 78% with On-Prem AI in 4 Weeks\",\"position\":3}]},{\"@id\":\"https:\/\/blog.forfis.com#org\",\"@type\":\"Organization\",\"name\":\"Forfis\",\"url\":\"https:\/\/blog.forfis.com\"}]}","geo_content_hash":"4e97c5fee9e6a304b0096aaf4860aafe7c242aee4c1761f3f0c9201bf2496a8e","footnotes":""},"categories":[37],"tags":[69,71,43],"class_list":["post-449","post","type-post","status-publish","format-standard","hentry","category-fintech-and-payments","tag-automate-monthly-reporting","tag-candidate-screening","tag-switzerland"],"_links":{"self":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/449","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/comments?post=449"}],"version-history":[{"count":0,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/449\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/media?parent=449"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/categories?post=449"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/tags?post=449"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}