TestCon Europe 2026

 

October 20-23

Vilnius & Online

Igor Dorovskikh

Founder

EnGenious

United States

About

Igor Dorovskikh is a QA and AI testing leader with over 15 years of experience building and scaling world-class test automation teams at companies such as Apple, Grammarly, and Tinder. He is the Founder of Engenious University and the CEO of Marathon Labs North America, where he helps engineers and organizations master AI testing, mobile automation, and CI/CD optimization. Igor is a frequent international speaker and is passionate about elevating QA professionals into strategic AI quality leaders.

Workshop

Igor Dorovskikh | Breaking GPT & Claude: AI Bug Bounty Workshop (Hands-On with Agenta.ai)

AI Testing, LLM Testing, Promptfoo, RedTeaming

1. Abstract
Hack AI Agents is a six-hour, hands-on workshop that turns QA engineers into AI-native testers. AI products fail in ways traditional test suites never catch - they hallucinate, leak bias, obey malicious prompts, and slip past safety guardrails. In this workshop you will learn how modern AI systems break, then test live LLMs for reliability, safety, and fairness using industry tooling like Promptfoo.
The format is a career accelerator built as a competitive hackathon. After a guided crash course in AI failure modes, teams move through timed challenge rounds - hunting hallucinations, exposing bias, running prompt-injection attacks, and probing safety guardrails - then document their findings and present them to the judges. You leave with reusable AI evaluation templates, documented findings, and the confidence to own AI quality end-to-end in AI-native products.

2. Agenda
9:00 – 9:20 Welcome & Challenge Brief
Event overview, goals and expected outcomes, judging criteria and scoring explained.
9:20 – 10:00 AI Testing Crash Course
The 7 categories of AI testing, why AI systems fail in production, and what makes a strong AI test case.
10:00 – 10:30 Promptfoo Walkthrough
Platform introduction, provided workflow overview, and what a successful evaluation run looks like.
10:30 – 10:55 Team Formation & Setup
Team creation, access check, first successful test run.
10:55 – 11:20 Break
11:20 – 12:10 Challenge Round 1: Hallucinations
Fact-checking outputs and identifying unsupported claims.
12:10 – 13:00 Challenge Round 2: Bias & Fairness
Testing for unequal behavior and comparing outputs across varied prompts.
13:00 – 14:00 Lunch Break
14:00 – 14:50 Challenge Round 3: Security
Prompt injection and instruction-override testing.
14:50 – 15:30 Challenge Round 4: Safety Guardrails
Probing refusal behavior, unsafe-content handling, and guardrail bypasses.
15:30 – 15:55 Break
15:55 – 16:35 Submission Preparation
Finalize findings, select top failures, write suggested fixes, prepare the final demo.
16:35 – 17:00 Presentations, Judging & Winners
Team demos, judge evaluation and feedback, awards and closing.

3. Objectives
By the end of the workshop, participants will be able to:
●    Recognize the seven categories of AI testing and explain why AI systems fail in production.
●    Design and run structured LLM evaluations with Promptfoo.
●    Detect hallucinations and identify unsupported or fabricated claims in model output.
●    Test for bias and fairness by comparing behavior across varied prompts and demographics.
●    Execute prompt-injection and instruction-override attacks to probe model security.
●    Assess safety guardrails and document guardrail bypasses.
●    Write clear, actionable findings with suggested fixes and present them to stakeholders.
Participants leave with reusable AI evaluation templates, documented findings, setup resources, and the identity shift into AI-native QA work.

4. Target Audience and Prerequisites
Who it's for: QA Engineers, SDETs, QA Leads, QA Managers, and technical testers who already understand quality and want to become fluent in AI reliability and evaluation. It suits people who want practical AI QA skills rather than generic AI inspiration, and who are ready to work in teams, compete, and present under pressure.
Not a fit for: those seeking a passive conference experience, pure beginners with no testing or technical background, or anyone who wants theory without implementation.
Prerequisites: No prior AI testing experience is required — the workshop opens with a guided introduction to AI failure modes and evaluation workflows before the challenge rounds begin. A basic testing or technical background is expected. Exposure to any programming language is helpful but not mandatory.

5. Technical Requirements
Tooling (walked through live):Must have Claude Code installed with minimum subscription.
Minimum system requirements:
●    macOS: macOS 13.5 (Ventura) or later (Apple Silicon), 8 GB RAM minimum (16 GB+ recommended), ~500 MB free disk space.
●    Windows: Windows 10 or 11 (64-bit), 8 GB RAM minimum (16 GB+ recommended), ~500 MB free disk space.
●    Stable internet connection for live LLM access.

2026-10-20

09:00

17:00