TestCon Europe 2026
![]()
October 20-23
![]()
Vilnius & Online

Susanne Pieterse
About
Susanne is an autodidact, full-stack software engineer, and iSAQB-certified software architect. She is passionate about creating innovative solutions and sharing knowledge. She also loves tea and heavy-bag boxing workouts.
Susanne Pieterse | Evaluating AI Agent Behaviour in RAG Systems
With LLMs and AI agents moving into production, teams quickly discover a hard truth: without structured evaluation, they cannot understand the impact of changes.
In this session, Susanne will show how to turn a retrieval-augmented generation (RAG)-based AI agent into a measurable system instead of a black box. Attendees will learn how to build a golden evaluation set and use it to continuously test the correctness, consistency, and source grounding of LLM outputs.
Susanne will demonstrate lightweight evaluation loops for regression testing prompts, models, and retrieval changes, making it possible to detect when updates silently degrade behaviour or introduce regressions.
She will also explore real-world failure modes such as hallucinations, broken retrieval, and behaviour drift, and how a simple evaluation setup can surface these issues before they reach production.
Takeaway: a practical approach to evaluating AI agents in RAG systems that enables safer changes, faster iteration, and production-grade confidence.