- Home
- Case Studies
- How We Fixed a Fintech AI Assistant That Ignored Its Knowledge Base
How We Fixed a Fintech AI Assistant That Ignored Its Knowledge Base
100%
Critical issues resolved
95%
Knowledge base accuracy achieved
70%
Reduction in QA time
Key Insights
Location
Germany
Project duration
6 weeks
Industry
Fintech
Technologies used
AWS Bedrock, Amazon Nova Pro, DeepEval, Amazon S3, PostgreSQL, AWS Lambda, API Gateway
Solutions
Table of Contents
- Project Overview
- Client Background and Existing AWS Setup
- The Challenge: Hallucinations and Ignored Knowledge Base Data
- Our Approach: Evaluation First, Then Model and Prompt Tuning
- Building an LLM Evaluation Pipeline with DeepEval
- Comparing Claude 3 Haiku, Claude 3.7 Sonnet and Amazon Nova Pro
- Post-Migration Issue Resolution
- Results: 95% Knowledge Base Accuracy in Six Weeks
- Key Quantitative Outcomes
- Impact Summary
- Conclusion & Next Steps
- FAQ
- Project Overview
- Client Background and Existing AWS Setup
- The Challenge: Hallucinations and Ignored Knowledge Base Data
- Our Approach: Evaluation First, Then Model and Prompt Tuning
- Building an LLM Evaluation Pipeline with DeepEval
- Comparing Claude 3 Haiku, Claude 3.7 Sonnet and Amazon Nova Pro
- Post-Migration Issue Resolution
- Results: 95% Knowledge Base Accuracy in Six Weeks
- Key Quantitative Outcomes
- Impact Summary
- Conclusion & Next Steps
- FAQ
Project Overview
A German B2B fintech company partnered with Perfsys to improve the reliability of its AWS-based AI assistant through AI model optimization. The client is a team of fewer than 10 people that runs a well-known platform of verified customer reviews for financial advisors, banks and insurers. The platform helps consumers make informed financial decisions.
The client's vision was to build a reliable, knowledge-based AI assistant that answers complex user queries from verified data and keeps context during long interactions.
Key takeaways
Perfsys built an automated DeepEval pipeline with 500+ test cases, moved the client's assistant from Claude 3 Haiku to Amazon Nova Pro and fixed four system-level issues. Within six weeks knowledge base accuracy reached 95%, every critical issue (including hallucination handling) was resolved and manual QA time dropped by 70%.
Client Background and Existing AWS Setup
The client had already implemented a serverless AWS architecture consisting of:
- Amazon Bedrock for AI inference
- Amazon S3 as a knowledge base repository
- AWS Lambda and API Gateway for orchestration
- A web UI for the frontend interface
Each AI agent represented a unique financial advisor persona sharing access to a centralized knowledge base stored in S3.

Despite this advanced setup, the agents were inconsistent, prone to hallucination and often ignored the knowledge base, which compromised reliability. The client engaged us to measure, diagnose and systematically improve agent performance.
The Challenge: Hallucinations and Ignored Knowledge Base Data
While the infrastructure was functional, the core challenge lay in AI quality and consistency:
- Agents forgot their personalities or initial instructions during extended conversations
- Context retention dropped significantly after 3 to 4 exchanges
- Agents produced hallucinated or incorrect answers, sometimes ignoring KB data
- No automated evaluation existed to track answer accuracy or reference validity
The client's main goal was clear:
"Ensure the AI agent provides accurate, reference-backed answers from the knowledge base, with measurable and repeatable quality metrics."
Our Approach: Evaluation First, Then Model and Prompt Tuning
Perfsys designed a three-phase strategy: automated evaluation first, then model experimentation, then prompt-level AI model optimization.
Building an LLM Evaluation Pipeline with DeepEval
We began by developing a custom Evaluation Pipeline based on the DeepEval framework. This pipeline allowed automatic testing of hundreds of AI interactions to measure:
- KB reference accuracy (how well answers use the knowledge base)
- Response consistency
- Invocation time (latency)
The evaluation pipeline enabled:
- Running 500+ automated test cases across multiple sessions
- Establishing quantitative baselines for each tested model
- Reproducing real user interaction patterns
This became the foundation for comparing every candidate model. If you are planning a similar build, our guide to building an AI agent MVP on Amazon Bedrock covers the architecture side.

Comparing Claude 3 Haiku, Claude 3.7 Sonnet and Amazon Nova Pro
We tested three models on Amazon Bedrock against the same evaluation set:
The testing revealed that Claude 3 Haiku, the client's initial choice, failed to reference the KB correctly in 80% of cases.
While Claude 3.7 Sonnet had better accuracy, Amazon Nova Pro offered optimal performance-to-cost ratio and superior consistency within Bedrock's ecosystem.
"The evaluation pipeline showed that Claude 3 Haiku, the client's original model, failed to reference the knowledge base in 80% of cases. Amazon Nova Pro gave us the best balance between accuracy and cost, and we resolved every critical issue within six weeks."
— Eugene Orlovsky, Founder and CEO, Perfsys

Post-Migration Issue Resolution
After moving to Amazon Nova Pro, we found and fixed four system-level issues:
Results: 95% Knowledge Base Accuracy in Six Weeks
Within six weeks, Perfsys successfully delivered a measurable improvement in AI performance and consistency through targeted AI model optimization.
Key Quantitative Outcomes
- 100% of critical issues resolved (language, fallback, KB consistency and hallucination handling)
- Knowledge base accuracy improved from 80% to 95%, ensuring nearly all answers are KB-based
- Evaluation automation reduced manual QA time by 70%, validating 500+ test cases per iteration
Impact Summary
- Valid answer consistency and reliability of KB usage significantly improved
- Invocation latency remained stable (~7 seconds average)
- Maintenance simplified through automated evaluation cycles
Conclusion & Next Steps
Through systematic testing and evaluation automation, plus Bedrock-native AI model optimization, we helped the client turn a poorly performing AI assistant into a reliable and measurable knowledge-based agent.
Next steps include:
- Expanding multilingual testing (DE, EN, FR)
- Integrating new agent personalities for domain-specific advisory roles
- Deploying the evaluation pipeline to monitor new model updates automatically
Update, October 2026: newer models such as Claude Sonnet 5.5 and Amazon Nova 2 have been released since this project. The evaluation pipeline can rerun the same test set against them, which is the reason we built it.
This kind of evaluation and optimization work is part of our AI agent development services on AWS Bedrock.

Perfsys builds and evaluates AI agents on Amazon Bedrock, and we can start by measuring what your assistant gets wrong.
FAQ
AI model optimization is the process of improving an AI model so it gives accurate and consistent answers in real-world use. It includes testing different models, refining prompts, tuning knowledge-base retrieval and measuring accuracy and latency. In short, it helps the AI use the right information, avoid hallucinations and run efficiently in production.
Knowledge base accuracy ensures the AI agent provides answers grounded in verified information instead of guessing. High KB accuracy improves answer quality and consistency, and it builds user trust, especially in industries like fintech.
Amazon Bedrock gives you one platform for comparing LLMs and connecting them to vector-based knowledge retrieval, plus tools for measuring performance. That makes evaluations repeatable, reduces hallucinations and tunes AI behavior without complex infrastructure.
The root cause was the model's weak use of the knowledge base. The initial Claude 3 Haiku model often generated answers from its own reasoning instead of referencing verified KB data stored in S3.
Knowledge base accuracy measures how closely an AI's answer aligns with the knowledge base. A high score means the response is grounded in approved information and free of invented content.
Amazon Nova Pro gave the best balance of cost and accuracy for this use case. Claude 3.7 Sonnet showed slightly higher accuracy, but Nova Pro integrated more efficiently with Amazon Bedrock and delivered more consistent performance at a better cost-to-quality ratio.
Perfsys implemented a repeatable evaluation pipeline, running 500+ automated test cases to compare model performance before and after changes. This let us track accuracy, latency and KB consistency with quantifiable metrics.
Amazon Bedrock: LLM hosting and knowledge base retrieval
DeepEval Framework: automated evaluation pipeline
Amazon S3 & PostgreSQL: vectorized knowledge-base storage
AWS Lambda + API Gateway: serverless orchestration
CEO & Founder @ Perfsys | Serverless architect with 10+ years of hands-on experience designing cloud-native architectures on AWS, backed by multiple AWS certifications. His writing bridges deep technical expertise with real-world business strategy, covering topics from AWS best practices to scaling tech-driven organizations.
Recommended for You
AWS Experts, On-Demand
Need to move fast? Our cloud team is ready to scale, secure, and optimize your systems. Get serverless expertise, 24/7 support, and seamless CI/CD pipelines when you need it most.
Please accept cookies to load the booking widget.
