Two MoTacon attendees are on the left. The MoTaacon logo is in the center, and to the right a prompt to Get Your Ticket.
Testing the untestable: building a strategy for testing AI thumbnail

Testing the untestable: building a strategy for testing AI

For decades, testing has relied on a simple truth: If I provide Input X, I should get Output Y.

Generative AI broke that truth. When you build AI Agents, the output changes every time. "Expected Results" do not exist in the same way. Traditional automation is brittle. Manual testing is too slow. And to make matters worse, results are subjective!

So, how do you assure quality in a system that, by its very nature, is unpredictable?

In this talk, we will outline the strategic concepts required to tame the chaos of testing AI by understanding the behaviour of the beast! To do this, we move beyond the code and explore the fundamental shifts in the Quality Lifecycle:

  • The Input Shift (Synthetic Personas): Moving from static test cases to Automated Persona-Driven Testing, using AI to simulate thousands of diverse user interactions (from "Upset existing customer" to "Confused new user").
  • The Verification Shift (LLM-as-a-Judge): Replacing binary assertions with Semantic Evaluation. We will discuss using models to grade "sentiment" and "safety" rather than just syntax.
  • The Baseline Shift (Benchmarking): How to establish a "Quality Baseline" for your product. We will cover how to measure if a new prompt is actually "better" or just "different" by tracking performance against a Golden Dataset.
  • The Safety Shift (HITL): Why automation isn't enough. We will discuss the role of Human-in-the-Loop (HITL) review and the unique value it adds to high-risk scenarios.
  • The Observability Shift (Tracing): AI is often a "black box." We discuss the importance of detailed tracing to understand not just what the model said, but why it decided to say it.

This is not a coding tutorial. This is a playbook for Quality Leaders who need to build a strategy for the next generation of software.

What you’ll learn
  • The Paradigm Shift: Understand the difference between Deterministic and Non-deterministic quality strategies, and why we must evolve beyond "Pass/Fail."
  • Core AI Testing Concepts: Gain a high-level understanding of Synthetic Personas, LLM-as-a-Judge, and Benchmarking—and where each fits in the delivery pipeline.
  • Managing Production Risk: Learn how Human-in-the-Loop (HITL) review captures the nuance and safety insights that automation cannot.

Comments

Sign in to comment
Explore MoT
From test cases to AI-ready quality data: why structure matters image
Learn how to structure and organize test management data as an AI-ready foundation that gives AI systems richer context for testing insights and workflows.
How to test GenAI Agents (by building one) image
Turn AI curiosity into engineering confidence.
Into The Motaverse image
Into the MoTaverse is a podcast by Ministry of Testing, hosted by Rosie Sherry, exploring the people, insights, and systems shaping quality in modern software teams.
Subscribe to our newsletter