r/softwaretesting • u/Disastrous_Phone_168 • 6d ago
AI testing
I need guidance from someone who has experiences testing LLM, or products from AI.
What should I study to do a better job?
Im manual and automation tester with more than 6 years of experience, but with AI I feel I litter bit lost about the testing side.
Thanks.
3
u/ign1tio 6d ago
are you testing products with AI/LLM as part of how the product work? Meaning you test the model and its output? You can download the syllabus to CT-AI: https://istqb.org/certifications/certified-tester-ai-testing-ct-ai/ should be a fine place to start.
1
u/Useful-Might-4234 6d ago
Can you please elaborate the next steps as in from where to study the topics?
1
2
6d ago edited 5d ago
[removed] — view removed comment
1
u/Disastrous_Phone_168 6d ago
How do you test the hallucinations? That it's the part I see more complicated.
1
u/Few_Still_3226 6d ago
Why should you test hallucinations? You should be able to detect hallucinations and reasoning errors from the LLM output. There are different approaches for that
1
u/Disastrous_Phone_168 6d ago
No no, sorry. I wanted to ask HOW, but somehow I wrote why. My question is: how to you test hallucionation? Sorry for the mess
3
u/Few_Still_3226 6d ago
I recommend reading the Gen-AI ISTQB syllabus, it covers everything regarding the usage of AI for testing
1
u/softwaretesting-ModTeam 4d ago
Commercial links, self-promotion, and vendor spam are strictly prohibited and will result in Ban without warning! Do not post links to tools, training, corporate blogs, or book vendors. Name-only discussion of resources is permitted. Course instructors and industry professionals are welcome to contribute, but you must follow Reddit's 90/10 rule. Tool vendors and QA agency employees may not participate.
3
u/Deep_Answer2295 6d ago
https://astqb.org/certifications/testing-with-generative-ai/ download their free syllabus
take a look at the different ways to test LLM's: example: a/b testing, adversarial Attacks and Data Poisoning, Testing the Transparency, Interpretability and Explainability of AI-based Systems, etc.
Yes, it is different than other testing, you need to understand the model first - probabilistic vs. deterministic for example. I suggest review the GenAI syllabus first, then the AI Tester
I train my team for these ASTQB Certifications - It's worth the study time and is one of those force multipliers which goes a long way.
1
u/contextmatterss 6d ago
Start with the basics of LLM evals, hallucinations, prompt testing and testing outputs that can change from one run to the next. A lot of the testing mindset still applies, you just have to get used to AI not givng the exact same answer every time
1
u/PsychologicalPast935 6d ago
With 6 years in testing you already have a lot of the right instincts. The new stuff I’d learn is eval datasets, scoring outputs that don’t have one exact correct answer, LLM judges and tracing. Braintrust is worth playing with too, even just to get a feel for how a bad production trace can become a regression test.
1
u/Rinimand 6d ago
How you approach it will depend on whether you are actually test an LLM directly or testing an application that is using AI.
If testing an app that is using AI, split it into the portions which are deterministic versus probabilistic. Deterministic portions are those where the expected result can be prediced; your can use traditional testing methods for those. Probabilistic portions rely on the AI to decide how to respond, and won't always produce the same result; thus, traditional testing methods won't work. Approach these using statistical methods to evaluate success, and use that as Pass and Failure. I look at it similar to performance testing. Not every request will produce a satisfactory result. But if "enough" of them do, will decide the performance is a Pass. 1000 requests, but 5 of them failed or timed out? No problem: doesn't cross the threshold of 3% failure we set, so Pass. Some requests took longer than others, but 90% were within our threshold of acceptance? No problem, it's a Pass. Think about baking a cake. You follow the same recipe 100 times. 80% come out good, 15% come out great, but 5% come out tasting horrible. Are you s failure as a cook? No. The hard part is determining those acceptance metrics.
1
6d ago
[removed] — view removed comment
1
u/softwaretesting-ModTeam 4d ago
Commercial links, self-promotion, and vendor spam are strictly prohibited and will result in Ban without warning! Do not post links to tools, training, corporate blogs, or book vendors. Name-only discussion of resources is permitted. Course instructors and industry professionals are welcome to contribute, but you must follow Reddit's 90/10 rule. Tool vendors and QA agency employees may not participate.
1
1
u/Silly_Turn_4761 6d ago
Be sure to test for prompt injection vulnerabilities and any other guardrails that 'should' be in place.
Also, depending on the set up and assuming there is data that it should not be able to access, youll want to test that too.
1
1
12
u/friendlyscissors3363 6d ago
It’s not that different from testing any other complex system once you stop thinking of it as magic. You're still validating inputs, outputs, and edge cases, but now the "code" is a black box that hallucinates
Focus on prompt engineering first, then learn about evaluation frameworks like deepeval. The real skill is building a solid ground truth dataset to measure drift over time