The Problem: Manual Invoice Queries Eating Analyst Hours
You run a 51-200 person fintech in Austria. Your finance and accounting team handles invoice processing, vendor reconciliation, and payment queries through SAP or Microsoft Dynamics ERP. Every week, a portion of your support tickets are routine: ‘What is the status of invoice INV-2024-0847?’, ‘Why was vendor X’s payment delayed?’, ‘What are the payment terms for this GL account?’ Each of these consumes 8-15 minutes of an analyst’s time, and the cost per ticket compounds across departments as you scale. The problem is not that your ERP is broken. It is that the knowledge needed to answer these questions is locked inside the ERP, and your team has to open the system, search, and interpret the data manually. A retrieval-augmented knowledge assistant built on pgvector embeddings search, integrated into your existing ERP via its API, can answer 60-75% of these queries without a human opening the system. The goal is not to replace your ERP. It is to lower the cost per support ticket by removing the manual search-and-interpret step from the workflow, while keeping a human in the loop for anything that touches money or a contract.
Prerequisites: What You Need Before Step 1
Before you start step 1, confirm the following are in place:
- ERP API access: You have read access to the SAP or Microsoft Dynamics ERP API for the invoice, vendor, and GL account objects. If you are on SAP S/4HANA, this means the OData API or the BAPI layer. If you are on Dynamics 365, this means the Web API or the OData endpoint. You do not need write access for the pilot.
- Invoice data in a queryable format: Your invoice records are stored in the ERP or in a connected document management system. PDFs are acceptable; the extraction step in the pilot will handle them.
- A measured baseline: You have logged the average cycle time and error rate for invoice-related support tickets over the last 30 days. This is your before/after reference. Without it, you cannot prove the pilot worked.
- A named pilot scope: One invoice-processing workflow, one department, one ERP instance. Do not attempt to cover all departments in the pilot.
- A human approver: A finance team member who will review any assistant output that touches a payment, a contract, or a vendor master data change. This person is part of the pilot, not an afterthought.
Step 1: Audit the Invoice Workflow and Pick the Pilot Scope
Run a process audit on your invoice-handling workflow. Map every step from invoice receipt to payment, and tag each step with the time it consumes and the error rate. For a typical Austrian fintech, the audit reveals that 40-60% of the cycle time is spent on data entry, status lookups, and reconciliation checks that do not require judgment. Identify the three to five workflows where the manual search-and-interpret step is the bottleneck. Document the ERP objects involved: which SAP tables or Dynamics entities hold the invoice, vendor, and GL account data. This audit output becomes the scope for the pilot. Do not skip this step. If you build the RAG assistant on the wrong workflow, the pilot will not reduce cost per ticket, and you will have spent a month on a system nobody uses.
Step 2: Build the pgvector Embeddings Schema
Design the pgvector schema that will store your invoice and ERP data as embeddings. Create a PostgreSQL table with a vector(1536) column (for OpenAI’s text-embedding-3-small) or vector(768) (for a local model like BGE-M3). Each row represents a chunk of invoice data: the invoice number, vendor name, GL account, amount, due date, and a short natural-language description of the transaction. For example, a row might look like: invoice_id: INV-2024-0847, vendor: 'Muster GmbH', gl_account: '4000', amount: 1250.00, due_date: '2024-09-15', description: 'Monthly SaaS subscription payment'. The description field is critical: it is what the LLM will use to ground its answer. Write it in plain language, not in ERP field codes. This step takes two to three days and is the foundation of the entire system.
Step 3: Ingest ERP Data and Generate Embeddings
Write the ingestion pipeline that pulls invoice and ERP data from SAP or Dynamics, extracts the relevant fields, generates the natural-language description, computes the embedding, and inserts the row into the pgvector table. For SAP, use the OData API or a BAPI call to read the invoice header and line items. For Dynamics, use the Web API. The pipeline runs on a schedule: nightly for new invoices, and on-demand when a finance team member triggers a re-index. The embedding model is called for each new chunk. If you are using OpenAI’s text-embedding-3-small, the cost is approximately $0.02 per 1,000 tokens, which is negligible for a 51-200 person firm. If you are using a local model on your own hardware, the cost is zero but the latency is higher. Log every ingestion run with a timestamp and a row count so you can audit the data flow later.
Step 4: Build the RAG Query Layer with Human-in-the-Loop Approval
Build the query interface that a finance team member will use. The user types a question in natural language, for example: ‘What is the status of invoice INV-2024-0847 and when is it due?’ The system embeds the question, runs a cosine-similarity search against the pgvector index, retrieves the top 5-8 chunks, and passes them as context to the LLM. The LLM is prompted to answer in the language of the query (German, English, or another supported language) and to cite the specific invoice number and GL account it is referencing. The response is displayed in a lightweight dashboard or integrated into your existing helpdesk. If the question involves a payment action, a vendor master data change, or a contract modification, the system flags it for human approval. The approver sees the assistant’s draft, the retrieved context, and a one-click approve or reject button. This step takes one to two weeks and is where the human-in-the-loop design becomes operational.
Step 5: Run the Pilot and Measure the Before/After Baseline
Run the pilot for four to six weeks on the single workflow you scoped in step 1. Measure the cycle time and error rate for every invoice-related ticket that passes through the assistant. Compare the numbers against your baseline from the prerequisites. The target is a 30-45% reduction in cycle time and a measurable drop in error rate. Track the escalation rate: how often does the assistant flag a query for human approval, and how often does the approver reject the assistant’s draft? If the escalation rate is above 20%, your retrieval thresholds are too loose or your natural-language descriptions in the pgvector table are too vague. Tune the top-k parameter and the similarity threshold. If the error rate does not drop, check whether the LLM is hallucinating invoice numbers or GL accounts that do not exist in the retrieved context. The pilot output is a one-page report with the before/after numbers, the escalation rate, and the list of queries that the assistant could not answer. This report is what you use to justify the rollout to additional departments.