TestCon Europe 2026

 

October 20-23

Vilnius & Online

Susanne Pieterse

Senior DevOps Engineer

EDSN

The Netherlands

About

Susanne is an autodidact, full-stack software engineer, and iSAQB-certified software architect. She is passionate about creating innovative solutions and sharing knowledge. She also loves tea and heavy-bag boxing workouts.

Talk

Susanne Pieterse | Evaluating AI Agent Behaviour in RAG Systems

AI Agent, RAG, Evaluation, GenAI, LLM

With LLMs and AI agents moving into production, teams quickly discover a hard truth: without structured evaluation, they cannot understand the impact of changes.

In this session, Susanne will show how to turn a retrieval-augmented generation (RAG)-based AI agent into a measurable system instead of a black box. Attendees will learn how to build a golden evaluation set and use it to continuously test the correctness, consistency, and source grounding of LLM outputs.

Susanne will demonstrate lightweight evaluation loops for regression testing prompts, models, and retrieval changes, making it possible to detect when updates silently degrade behaviour or introduce regressions.

She will also explore real-world failure modes such as hallucinations, broken retrieval, and behaviour drift, and how a simple evaluation setup can surface these issues before they reach production.

Takeaway: a practical approach to evaluating AI agents in RAG systems that enables safer changes, faster iteration, and production-grade confidence.

2026-10-22

10:10

10:55

Hall 3