Perfsys — AWS Consulting Partner

How We Fixed a Fintech AI Assistant That Ignored Its Knowledge Base

100%

Critical issues resolved

95%

Knowledge base accuracy achieved

70%

Reduction in QA time

Key Insights

Location

Germany

Project duration

6 weeks

Industry

Fintech

Technologies used

AWS Bedrock, Amazon Nova Pro, DeepEval, Amazon S3, PostgreSQL, AWS Lambda, API Gateway

Solutions

Project Overview

A German B2B fintech company partnered with Perfsys to improve the reliability of its AWS-based AI assistant through AI model optimization. The client is a team of fewer than 10 people that runs a well-known platform of verified customer reviews for financial advisors, banks and insurers. The platform helps consumers make informed financial decisions.

The client's vision was to build a reliable, knowledge-based AI assistant that answers complex user queries from verified data and keeps context during long interactions.

Key takeaways

Perfsys built an automated DeepEval pipeline with 500+ test cases, moved the client's assistant from Claude 3 Haiku to Amazon Nova Pro and fixed four system-level issues. Within six weeks knowledge base accuracy reached 95%, every critical issue (including hallucination handling) was resolved and manual QA time dropped by 70%.

Client Background and Existing AWS Setup

The client had already implemented a serverless AWS architecture consisting of:

  • Amazon Bedrock for AI inference
  • Amazon S3 as a knowledge base repository
  • AWS Lambda and API Gateway for orchestration
  • A web UI for the frontend interface

Each AI agent represented a unique financial advisor persona sharing access to a centralized knowledge base stored in S3.

Initial AWS serverless architecture of the fintech AI assistant: Bedrock agents, S3 knowledge bases, Lambda and API Gateway
Initial AWS serverless architecture used by the client, including S3-based knowledge banks, Bedrock LLM agents, Lambda orchestration and API Gateway endpoints.

Despite this advanced setup, the agents were inconsistent, prone to hallucination and often ignored the knowledge base, which compromised reliability. The client engaged us to measure, diagnose and systematically improve agent performance.

The Challenge: Hallucinations and Ignored Knowledge Base Data

While the infrastructure was functional, the core challenge lay in AI quality and consistency:

  • Agents forgot their personalities or initial instructions during extended conversations
  • Context retention dropped significantly after 3 to 4 exchanges
  • Agents produced hallucinated or incorrect answers, sometimes ignoring KB data
  • No automated evaluation existed to track answer accuracy or reference validity

The client's main goal was clear:

"Ensure the AI agent provides accurate, reference-backed answers from the knowledge base, with measurable and repeatable quality metrics."

Our Approach: Evaluation First, Then Model and Prompt Tuning

Perfsys designed a three-phase strategy: automated evaluation first, then model experimentation, then prompt-level AI model optimization.

Building an LLM Evaluation Pipeline with DeepEval

We began by developing a custom Evaluation Pipeline based on the DeepEval framework. This pipeline allowed automatic testing of hundreds of AI interactions to measure:

  • KB reference accuracy (how well answers use the knowledge base)
  • Response consistency
  • Invocation time (latency)

The evaluation pipeline enabled:

  • Running 500+ automated test cases across multiple sessions
  • Establishing quantitative baselines for each tested model
  • Reproducing real user interaction patterns

This became the foundation for comparing every candidate model. If you are planning a similar build, our guide to building an AI agent MVP on Amazon Bedrock covers the architecture side.

Multi-phase evaluation timeline for the Bedrock AI assistant, ending with the move to Amazon Nova Pro
Multi-phase agent evaluation journey showing how Perfsys refined the testing pipeline and validated model performance before selecting Amazon Nova Pro.

Comparing Claude 3 Haiku, Claude 3.7 Sonnet and Amazon Nova Pro

We tested three models on Amazon Bedrock against the same evaluation set:

Model
Claude 3 Haiku
Claude 3.7 Sonnet
Amazon Nova Pro
KB Reference Failure Rate
80% failure
15% failure
19% failure
Invocation Time
5.6 sec
6.8 sec
7.7 sec
Cost/Performance Notes
Fast, but unreliable KB referencing
Accurate, but higher cost
Best balance of speed, cost, and accuracy

The testing revealed that Claude 3 Haiku, the client's initial choice, failed to reference the KB correctly in 80% of cases.

While Claude 3.7 Sonnet had better accuracy, Amazon Nova Pro offered optimal performance-to-cost ratio and superior consistency within Bedrock's ecosystem.

"The evaluation pipeline showed that Claude 3 Haiku, the client's original model, failed to reference the knowledge base in 80% of cases. Amazon Nova Pro gave us the best balance between accuracy and cost, and we resolved every critical issue within six weeks."

— Eugene Orlovsky, Founder and CEO, Perfsys

Model comparison results showing knowledge base reference failure rates for Claude 3 Haiku, Claude 3.7 Sonnet and Amazon Nova Pro
Model comparison results showing KB reference failure rates for Claude 3 Haiku, Claude 3.7 Sonnet and Amazon Nova Pro.

Post-Migration Issue Resolution

After moving to Amazon Nova Pro, we found and fixed four system-level issues:

Issue
Language Support
Contact Info Retrieval
KB File Conflicts
Out-of-Scope Handling
Description
Agent defaulted to German only
Failed to provide consultant details
Duplicate answers in S3 files
Agents generated irrelevant or invented answers
Resolution
Updated system prompt
Added structured fallback prompts
Logic updated to prefer most recent version
Improved fallback strategy
Status
Resolved
Resolved
Resolved
Resolved

Results: 95% Knowledge Base Accuracy in Six Weeks

Within six weeks, Perfsys successfully delivered a measurable improvement in AI performance and consistency through targeted AI model optimization.

Key Quantitative Outcomes

  • 100% of critical issues resolved (language, fallback, KB consistency and hallucination handling)
  • Knowledge base accuracy improved from 80% to 95%, ensuring nearly all answers are KB-based
  • Evaluation automation reduced manual QA time by 70%, validating 500+ test cases per iteration

Impact Summary

  • Valid answer consistency and reliability of KB usage significantly improved
  • Invocation latency remained stable (~7 seconds average)
  • Maintenance simplified through automated evaluation cycles

Conclusion & Next Steps

Through systematic testing and evaluation automation, plus Bedrock-native AI model optimization, we helped the client turn a poorly performing AI assistant into a reliable and measurable knowledge-based agent.

Next steps include:

  • Expanding multilingual testing (DE, EN, FR)
  • Integrating new agent personalities for domain-specific advisory roles
  • Deploying the evaluation pipeline to monitor new model updates automatically

Update, October 2026: newer models such as Claude Sonnet 5.5 and Amazon Nova 2 have been released since this project. The evaluation pipeline can rerun the same test set against them, which is the reason we built it.

This kind of evaluation and optimization work is part of our AI agent development services on AWS Bedrock.

Is your AI assistant still ignoring its knowledge base?
Is your AI assistant still ignoring its knowledge base?

Perfsys builds and evaluates AI agents on Amazon Bedrock, and we can start by measuring what your assistant gets wrong.

Book a Discovery Call
AWS AI Agents

FAQ

AI model optimization is the process of improving an AI model so it gives accurate and consistent answers in real-world use. It includes testing different models, refining prompts, tuning knowledge-base retrieval and measuring accuracy and latency. In short, it helps the AI use the right information, avoid hallucinations and run efficiently in production.

Knowledge base accuracy ensures the AI agent provides answers grounded in verified information instead of guessing. High KB accuracy improves answer quality and consistency, and it builds user trust, especially in industries like fintech.

Amazon Bedrock gives you one platform for comparing LLMs and connecting them to vector-based knowledge retrieval, plus tools for measuring performance. That makes evaluations repeatable, reduces hallucinations and tunes AI behavior without complex infrastructure.

The root cause was the model's weak use of the knowledge base. The initial Claude 3 Haiku model often generated answers from its own reasoning instead of referencing verified KB data stored in S3.

Knowledge base accuracy measures how closely an AI's answer aligns with the knowledge base. A high score means the response is grounded in approved information and free of invented content.

Amazon Nova Pro gave the best balance of cost and accuracy for this use case. Claude 3.7 Sonnet showed slightly higher accuracy, but Nova Pro integrated more efficiently with Amazon Bedrock and delivered more consistent performance at a better cost-to-quality ratio.

Perfsys implemented a repeatable evaluation pipeline, running 500+ automated test cases to compare model performance before and after changes. This let us track accuracy, latency and KB consistency with quantifiable metrics.

Amazon Bedrock: LLM hosting and knowledge base retrieval

DeepEval Framework: automated evaluation pipeline

Amazon S3 & PostgreSQL: vectorized knowledge-base storage

AWS Lambda + API Gateway: serverless orchestration

Eugene Orlovsky
Eugene Orlovsky

CEO & Founder @ Perfsys | Serverless architect with 10+ years of hands-on experience designing cloud-native architectures on AWS, backed by multiple AWS certifications. His writing bridges deep technical expertise with real-world business strategy, covering topics from AWS best practices to scaling tech-driven organizations.

Recommended for You

View All News

AWS Experts, On-Demand

Need to move fast? Our cloud team is ready to scale, secure, and optimize your systems. Get serverless expertise, 24/7 support, and seamless CI/CD pipelines when you need it most.

Please accept cookies to load the booking widget.