The Problem: Senior Staff Buried in Routine Queries
A 12-person fintech in Germany runs on senior engineers and compliance officers who spend 30-40% of their week answering the same questions: “What is our KYC threshold for a new merchant?” “How do we process a chargeback for a card issued in 2019?” “Where is the latest version of our AML policy?” The answers live in Notion, Confluence, and a helpdesk that no one has reorganized since the last product launch. Every query pulls a senior person off their actual work. The cost is not just time—it is the compounding drag on a team that cannot hire a dedicated support layer because the headcount budget is already committed to product and compliance.
The fix is not a chatbot bolted onto a Slack channel. It is a retrieval-augmented generation (RAG) pipeline that ingests the existing documentation, a predictive scoring model that routes incoming tickets by risk, and a human-in-the-loop approval layer that keeps money-touching actions under human control. The architecture is model-agnostic: OpenAI or Anthropic APIs where quality matters, open-weight models on the client’s own hardware where regulated data cannot leave the building. The integration point is the helpdesk and the documentation platform—Notion or Confluence—via their existing APIs. No new SaaS stack. No rip-and-replace.
Mechanism: RAG Pipeline and Predictive Scoring
The pipeline has three stages: ingestion, retrieval, and generation.
Ingestion. The system pulls documents from Notion or Confluence via their REST APIs. Each document is chunked into 256-512 token segments using a sliding window with 50-token overlap. A sentence-transformer model—BGE-M3 or OpenAI’s text-embedding-3-small—converts each chunk into a 1024-dimensional vector. These vectors store in pgvector, a PostgreSQL extension that adds cosine-similarity search to a standard Postgres instance. For a 10,000-document corpus, the initial index build takes under 5 minutes on a single VPS with 16 GB RAM.
Retrieval. When a user types a query, the same embedding model converts it to a vector. pgvector returns the top-k (typically k=5) most similar chunks using cosine distance. The query is augmented with metadata filters—document type, last-updated date, access level—so the retrieval respects the team’s existing permission model.
Generation. The retrieved chunks, the original query, and a system prompt feed into an LLM. The model generates an answer grounded in the retrieved text, with inline citations pointing to the source document and section. For a fintech, the system prompt explicitly instructs the model to flag any answer that touches payment thresholds, AML rules, or contract terms for human review before it reaches the user.
The predictive scoring model runs in parallel. It is a lightweight classifier—logistic regression or a small feedforward network—trained on historical helpdesk tickets. Features include sender email domain, ticket subject keywords, document type referenced, and time-of-day. The output is a probability score: P(fraud-related), P(AML-related), P(routine). Tickets scoring above 0.7 on fraud or AML route directly to a senior compliance officer. Lower-scoring tickets get an AI-drafted first response for human approval in the helpdesk queue.
Trade-offs: Model Choice, Chunking, and Approval Scope
The architect faces three major trade-offs, each with a concrete cost.
Model choice: cloud API vs. on-premises. OpenAI’s gpt-4o or Anthropic’s claude-3-5-sonnet deliver higher answer quality than open-weight models like Llama 3 70B or Mistral 8x7B. But for a German fintech handling payment data, sending customer names and transaction details to a US-based API may violate internal data-residency policies. The cost of going on-premises: you need a GPU with at least 24 GB VRAM (an A100 or a used RTX 4090 cluster), and the model’s answer quality drops by 10-15% on complex multi-step queries. The mitigation is hybrid: use cloud APIs for internal documentation queries where no customer data is involved, and open-weight models for anything that touches customer PII or payment records.
Chunking strategy: fixed-size vs. semantic. Fixed 512-token chunks are simple and fast. Semantic chunking—splitting on paragraph boundaries, headings, or natural language breaks—improves retrieval precision by 8-12% but adds complexity to the ingestion pipeline. For a 12-person team, fixed-size chunking with 50-token overlap is the pragmatic default. Semantic chunking becomes worth the engineering time once the corpus exceeds 50,000 documents.
Human-in-the-loop scope: all responses vs. risk-based. Requiring human approval for every AI-generated response defeats the purpose of automation. The risk-based approach—approve only responses touching money, health data, or contracts—reduces the approval queue by 60-70% while keeping regulatory accountability. The cost: you must define the risk categories precisely and build the routing logic into the helpdesk workflow. For a fintech, the categories are clear: payment processing, AML/KYC, contract terms, and anything involving a customer’s financial data.
Recommendation: 8-Week Pilot Scope for a 12-Person Fintech
For a 12-person fintech in Germany, the 8-week pilot follows a fixed scope: one process, one data source, one measurable outcome.
Weeks 1-2: Process audit. Map the current workflow. Measure baseline cycle time for internal knowledge queries (target: 15-20 minutes per query) and ticket triage error rate (target: 10-15% misclassification). Identify the single highest-ROI process—usually internal knowledge search or ticket triage. Confirm the data source: Notion, Confluence, or both. Document the permission model so the RAG pipeline respects access levels.
Weeks 3-5: Build. Ingest the documentation corpus into pgvector. Train the predictive scoring model on 6-12 months of historical helpdesk tickets. Build the RAG pipeline with the chosen LLM backend. Integrate with the helpdesk via its API so AI-drafted responses appear in the agent’s queue with confidence scores and source citations.
Weeks 6-7: Integration and UAT. Connect the pipeline to Notion/Confluence for real-time document updates. Run user acceptance testing with 3-5 senior staff. Measure cycle time and error rate against the baseline. Adjust the risk-based approval thresholds based on UAT feedback.
Week 8: Go-live and baseline report. Ship the pilot. Produce a before/after report showing cycle time reduction (target: 15-20 min → under 2 min) and error rate change (target: 30-50% reduction in misclassification). The report becomes the business case for rollout to additional processes in subsequent 4-6 week sprints.
The architecture is deliberately model-agnostic. If the team later migrates from OpenAI to Anthropic, or from cloud to on-premises, the RAG pipeline, embedding model, and scoring logic remain unchanged. The integration point is the LLM API call, not the entire stack.
Leave a Reply