Ujjwal Kumar Singh | Comments
With servers in >250 cities around the world, check your site for localization problems, broken GDPR banners, etc.
Congrats Gary...
Wow...
It feels great to be part of something amazing.
It was really good.
On the SME dependency question, testers can definitely help design golden scenarios, but they bring a different perspective. SMEs know what the correct outcome should be, while testers focus on finding situations where the system may confidently produce the wrong answer. Both roles are important and complement each other.
Many golden datasets fail because they focus only on happy paths instead of the failure cases that matter most.
As for the 80% metric, it only means something if you define what the 80% represents. A pass rate across 50 similar scenarios tells you very little. What matters is coverage across different failure patterns.
In LLM evaluation, these failures are often different from traditional testing. From my experience with MCP agent workflows, common issues include context carrying over between steps, unclear tool descriptions causing wrong routing, and responses that look correct but fail on edge cases.
I think it should change. Just because something is generated by AI doesn't mean it should be considered done. It still needs to be evaluated against the same quality standards we expect from human-written work.
Today, AI generates code, AI reviews code, and in many cases AI even suggests fixes. If there is no meaningful human evaluation in the loop, who is accountable for confirming that the work actually meets the required quality?
AI has become a crucial part of software development, but that also means our definition of Done should evolve. Just because AI completed it shouldn't automatically mean it's ready.
There are many open source communities where testers can contribute, including Selenium, Playwright, Appium, and Linux.
Most support contributors through documentation, issue trackers, discussions, and code reviews.
Linux has one of the largest communities, but its complexity can be overwhelming for beginners. Projects like Selenium, Playwright, and Appium are often easier starting points.
With maintainers handling a growing number of pull requests, contributors are increasingly expected to learn through documentation and contribution guides.
The struggle of breaking the stereotype that testing brings quality is one of the biggest challenges I have faced.
Another challenge for me is convincing the team in priortizing test execution over test cases. Documentation are good but minimum documentation can help in saving time for better test strategy, test design, etc.
Guardian of the Galaxy - X
Guardian of the Motaverse - ✓
Glad to be part of MoTAVERSE as an ambassador.
Ujjwal Kumar Singh
SDET @ Skeps
He/Him
Hi, I’m Ujjwal, a software tester and quality advocate. Exploring how quality works beyond tools and into systems, decisions, and trade-offs.
Substack: https://substack.com/@beinghumantester
Open To
Speak
Write
Podcasting
Teach
Work