Two MoTacon attendees are on the left. The MoTaacon logo is in the center, and to the right a prompt to Get Your Ticket.

Goldens for your evals

25 Jun 2026

In this moment: Swathika Visagn
I am exploring how to evaluate AI agent workflow and just learnt the golden test cases that we provide to any 'LLM evaluation f/w' is crucial.
From a traditional QA mindset, I would think of positive, negative and edge case scenarios. From an AI evaluation mindset, we need to look into this in a data science lens. 

I am thinking what would be a convincing range of data or scenarios that I can feed to the evaluation framework? 
How to shift left and prepare these golden test cases in advance by working closely with SMEs ?
Do these golden test cases be part of a requirement gathering activity than waiting till testing ?
Does it sound meaningful, if I say I evaluated 50 scenarios and the metric score was 80% ?
Can testers come up with 50 different scenarios without being an SME of the domain where the agent will be used ?

Any thoughts and tips ?
Swathika Visagn profile image
Swathika Visagn
Senior Test Engineer at PwC UK

I am a very curious Senior Quality Engineer who is more driven towards automation and promotes shifting left. I have proven experience as an agile tester having strong fundamentals in manual and automation testing principles. I enjoy the entire journey from setting the automation framework from scratch to building the pipelines onto continuous integration tools like Jenkins.

My framework adds more flavor by incorporating service layer (APIs) calls with UI layer automation which we call it 'Seaming' in automation terms. I communicate with stakeholders about risks, accessibility and pain points rather than number of passes/fails.

I test with a purpose by automating business flow and add in appropriate plug-ins to make the automation reports/metrics readable for stakeholders. I also love to take part in agile ceremonies and volunteer to run retrospectives/daily scrums to keep the team self thriving in temporary absence of the scrum master.

When I'm not scripting, I love to binge on Netflix, indulge in testing communities, read about Web3 and all things Quality :-) I am a yogic person too. If anything that calms me that's a cup of chai and a morning walk in the park.

Open To
Write
Team Account Member
Attending MoTaCon 🤝
Jas Manigundan
Hi Swathika, sharing some context from my experience over the last year. Hope it helps. - Evaluation Scenarios - Start small with a "golden dataset" (1–3 examples per scenario). It aligns the team on "what good looks like" without becoming a maintenance burden. - Shifting Left: Integrate test-case preparation directly into user stories and acceptance criteria. Leaving this until after launch is incredibly painful and doesn't scale. The whole team should own this and contribute. - Golden data as part of Requirements - 100% yes. In an indeterministic LLM world, building before defining your expected responses makes no sense. - Metrics (50 scenarios / 80%) - Numbers alone are meaningless. An 80% score is a failure if the missing 20% covers your core features. Assign weightage to critical scenarios, start small, and iterate. It depends on the domain you are in - for example, bias might have more weightage than another factor for certain industries like HR.

Ujjwal Kumar Singh
On the SME dependency question, testers can definitely help design golden scenarios, but they bring a different perspective. SMEs know what the correct outcome should be, while testers focus on finding situations where the system may confidently produce the wrong answer. Both roles are important and complement each other. Many golden datasets fail because they focus only on happy paths instead of the failure cases that matter most. As for the 80% metric, it only means something if you define what the 80% represents. A pass rate across 50 similar scenarios tells you very little. What matters is coverage across different failure patterns. In LLM evaluation, these failures are often different from traditional testing. From my experience with MCP agent workflows, common issues include context carrying over between steps, unclear tool descriptions causing wrong routing, and responses that look correct but fail on edge cases.

Swathika Visagn
great inputs thanks all :)

Sign in to comment
Explore MoT
Influence, from the other side of the table image
What I learned about influence by becoming a stakeholder
MoT Software Testing Essentials Certificate image
Boost your career in software testing with the MoT Software Testing Essentials Certificate. Learn essential skills, from basic testing techniques to advanced risk analysis, crafted by industry experts.
This Week in Quality image
Debrief the week in Quality via a community radio show hosted by Simon Tomes and members of the community
Subscribe to our newsletter