Two MoTacon attendees are on the left. The MoTaacon logo is in the center, and to the right a prompt to Get Your Ticket.
Test Oracle image
  • Demi Van Malcot's profile image
A reference point used to decide whether a test has passed or failed. For LLM testing this becomes unreliable because multiple valid outputs can exist for the same input. So what? The absence of a stable oracle is one of the central challenges of AI testing. Techniques like metamorphic testing exist partly to work around it by checking consistency rather than correctness. Example: "Who was the first president of the USA?" has a clear oracle. A summarisation request does not. "The 'expected result' can be determined, but will always have some ambiguity. Comparing it to the 'actual result' won't be as straightforward as you are used to." — Demi Van Malcot
Deep Learning image
  • Demi Van Malcot's profile image
A machine learning technique that uses layered neural networks to find patterns across large volumes of data. LLMs are built on deep learning to make connections across billions of words and generate contextually relevant responses. So what? LLMs work by statistical pattern-matching rather than reasoning — a foundational insight for anyone designing tests. Example: An LLM predicts the most statistically likely next word or phrase, not the most factually accurate one. "They are trained on billions of words from different sources. Using deep learning, they make connections between all the words they are trained on to answer whatever questions we ask of them." — Demi Van Malcot
Generative AI image
  • Demi Van Malcot's profile image
A category of AI model that produces new content: text, images, code, or other outputs, in response to input prompts, rather than returning a fixed or pre-programmed answer. LLMs are the most widely used type of generative AI. So what? Because outputs are generated fresh each time, the testing approaches used for traditional deterministic software don't transfer cleanly. Concepts like "expected result" and pass/fail need to be rethought. Example: ChatGPT, Claude, Bard, and Copilot are all generative AI applications built on large language models. "It won't give an answer based on what is logically correct, but on what is statistically most likely. The sentences can be completely correct, while the answer is completely wrong." — Demi Van Malcot
A practical introduction to testing LLMs image
  • Demi Van Malcot's profile image
Learn how to evaluate LLM quality and limitations using a range of testing techniques, from unit and regression testing to bias, adversarial and explainability testing.
Everytime I have to confirm the information is real and correct when filling in testdata image
We have recreated the import screens users need to fill in for declaration in our test environment. Although I appreciate the devs wanting to make it as close to real as possible, I have to check t...
Changing the conversation changes the future image
  • Abby Bangser's profile image
  • Simon Tomes's profile image
  • Gary Hawkes's profile image
  • Aj Wilson's profile image
  • Demi Van Malcot's profile image
  • TWiQ — This Week in Quality's profile image
In TWiQ today, Aj raised the idea of self-service infrastructure and it got me thinking about how that is related to platform engineering. Something I realise I don't know enough about, but a recen...
AMA about how baking a cake is the same as developping software image
Inspired by my manager who explained I 'monkey-barred' to a new job (before I was a tester I was a baker) and couldn't really transfer a lot of skills with me. I looked at him baffled because not o...
TWiQ - Episode 139 -  Cameras On and Our Hosts flying in from Left and Right image
  • Simon Tomes's profile image
  • Demi Van Malcot's profile image
  • TWiQ — This Week in Quality's profile image
TWiQ episodes now sounds more cool when we see our hosts with cameras on and their entry reminds me of animations - Flying In from Left and Right :) Cool discussions about negative tests, happ...
Future platforming conversations, right here, right now - TWIQ Ep 139 image
  • Simon Tomes's profile image
  • Rosie Sherry's profile image
  • Gary Hawkes's profile image
  • Demi Van Malcot's profile image
Live experimentation, technical glitches, and community challenges along the way
Can you see me now?  image
  • Simon Tomes's profile image
  • Demi Van Malcot's profile image
Today's This Week in Quality had the hosts with their cameras on for the very first time! ... And since I could see them, I was extremely convinced that they could see me as well. Cue lots of emoti...
High value activities: The faces of quality image
  • Simon Tomes's profile image
  • Gary Hawkes's profile image
  • Preeti Gupta's profile image
  • Demi Van Malcot's profile image
  • TWiQ — This Week in Quality's profile image
A This Week in Quality (TWiQ) evolution, we switched on the cameras for the first time! Thank you to Demi, Preeti and Gary for being the first people to do that. 🎉Do listen and now watch the episod...
TWiQ, camera, action: Vibe mapping high value activities - TWIQ Ep 138 image
  • Simon Tomes's profile image
  • Gary Hawkes's profile image
  • Preeti Gupta's profile image
  • Demi Van Malcot's profile image
This Week in Quality enters a new, cameras-on format
Subscribe to our newsletter