AI Risk Deck: Discrimination & Toxicity
27 Aug 2026
The overwhelm of AI and risks associated with it is real, so I thought I'd have a go at pulling some data together to create lists and helpful tools to help us all explore AI Risks better.
The AI Risk Deck for Discrimination & Toxicity
What is it? A list of risks related to AI and Discrimination & Toxcity.
Where does the data come from? I sourced it from the MIT AI Risk Initiative and asked Claude to research and structure it.
What can you do with the list? Bookmark it. Refer to it. Share it with your team. Use it for exploratory testing guidance. With the help of AI tools, build internal tools with it.
AI Risk Deck: Discrimination & Toxicity. On SpaceDucking (a website we've put together to enable practice testing and host little tools), I built a visual deck as an example of how to use this data in a different and helpful way to enable you and your team to think about how to test and implement quality practices. Play with the AI Risk Deck: Discrimination & Toxicity.
1. Unfair discrimination and misrepresentation
Biased training data
Historical and societal inequalities present in the data before the model is ever run.
Historical and societal inequalities present in the data before the model is ever run.
Sub-categories: unrepresentative sampling; over-weighted and under-weighted groups; English and Western-dominant corpora; misogynistic, ageist and white supremacist content over-represented in scraped internet text; uneven pronoun and identity prevalence; data contamination; synthetic data trained on earlier biased models.
Model and algorithmic bias
Bias introduced by choices that have nothing to do with the data.
Sub-categories: model architecture selection; regularisation and optimisation technique; presentation bias; model evaluation bias; popularity bias; compression techniques and hardware choices that amplify harm on protected attributes; unintentional bias amplification, where the output is more biased than the dataset it came from.
Discriminatory decisions and allocative harm
Withholding opportunities, resources or entitlements from particular groups.
Sub-categories: opportunity loss; economic loss; benefits and entitlements loss; hiring, lending, housing, welfare, healthcare, education and law enforcement contexts; protected characteristics; the proxy problem, where stripping race and gender from training data fails because models infer them from names, locations and other apparently unrelated fields.
Stereotyping and representational harm
Reproducing unjust social hierarchies through how groups are depicted, categorised or left out.
Reproducing unjust social hierarchies through how groups are depicted, categorised or left out.
Sub-categories: stereotyping social groups; demeaning social groups; erasing social groups; alienating social groups; denying people the chance to self-identify; reifying essentialist categories; homogenisation; misrepresentation, over-representation and under-representation; dignity loss.
Exclusionary norms
Language carries social categories, and a model that faithfully encodes language encodes who those categories leave out.
Language carries social categories, and a model that faithfully encodes language encodes who those categories leave out.
Sub-categories: definitional exclusion (e.g. "family" as married opposite-sex parents with a blood-related child); implied gender or ethnic identity in assistant design (gendered names, vernacular, product descriptions); monolingual bias; cultural dispossession; erasure of ways of speaking, humour and voice that carry cultural identity.
Preference, ideological and political bias
The model's view of contested questions, presented as neutral.
The model's view of contested questions, presented as neutral.
Sub-categories: political leanings in output; value embedding, where developers pick normative values in the absence of agreed standards; value lock-in from models never retrained as society shifts; outcome homogenisation from many deployers using the same foundation model; ideological homogenisation at global scale; sycophancy, reporting the user-preferred answer rather than the correct one.
Bias inside moderation and safety systems
The guardrails discriminate too, and the people they fail are the people they were meant to protect.
Sub-categories: shadowbanning and disproportionate suppression; dialect-insensitive toxicity detection flagging minority speech as offensive; community erasure; marginalised groups paying more of the cost of an intervention they benefit from less; automated removal performing close to randomly for some populations.
The guardrails discriminate too, and the people they fail are the people they were meant to protect.
Sub-categories: shadowbanning and disproportionate suppression; dialect-insensitive toxicity detection flagging minority speech as offensive; community erasure; marginalised groups paying more of the cost of an intervention they benefit from less; automated removal performing close to randomly for some populations.
Undetectable and unchallengeable bias
Discrimination that cannot be proven, contested or fixed.
Sub-categories: black-box models; explainability techniques that fail to surface discriminatory bias; manipulated explanations that hide sensitive attributes and substitute acceptable ones; unrepresentative risk testing; lack of testing diversity; opacity blocking liability and redress.
Discrimination that cannot be proven, contested or fixed.
Sub-categories: black-box models; explainability techniques that fail to surface discriminatory bias; manipulated explanations that hide sensitive attributes and substitute acceptable ones; unrepresentative risk testing; lack of testing diversity; opacity blocking liability and redress.
2. Exposure to toxic content
Toxic and offensive language
Rude, disrespectful or hostile output aimed at a person or group.
Sub-categories: profanity; insults and personal attacks; threats; identity attacks; slurs; cursing; scorn and impoliteness; implicit toxicity carried through sarcasm, irony and humour.
Rude, disrespectful or hostile output aimed at a person or group.
Sub-categories: profanity; insults and personal attacks; threats; identity attacks; slurs; cursing; scorn and impoliteness; implicit toxicity carried through sarcasm, irony and humour.
Hate speech and dehumanisation
Content that demeans or dehumanises people on the basis of sensitive personal characteristics.
Sub-categories: inciting, promoting or expressing hatred; demeaning and derogatory remarks about mental capacity, sensory and physical attributes, and behavioural attributes; perpetuating harmful beliefs; exclusion and isolation, whether social, political or economic.
Content that demeans or dehumanises people on the basis of sensitive personal characteristics.
Sub-categories: inciting, promoting or expressing hatred; demeaning and derogatory remarks about mental capacity, sensory and physical attributes, and behavioural attributes; perpetuating harmful beliefs; exclusion and isolation, whether social, political or economic.
Harassment and abuse
Behaviour that leaves an individual or group feeling alarmed or threatened.
Sub-categories: bullying; intimidation; shaming; humiliation; provoking; trolling; doxxing; sexual harassment; emotional abuse.
Behaviour that leaves an individual or group feeling alarmed or threatened.
Sub-categories: bullying; intimidation; shaming; humiliation; provoking; trolling; doxxing; sexual harassment; emotional abuse.
Violence and extremism
Content that promotes, glorifies, depicts or provides support for violence and extremist causes.
Content that promotes, glorifies, depicts or provides support for violence and extremist causes.
Sub-categories: supporting malicious organised groups; celebrating suffering; describing or endorsing violent acts; depicting violence and gore; weapon usage and development; military and warfare content; incitement to mass violence and genocide; terror content.
Sexual content
Explicit sexual material generated or shown to people who did not consent to it or should not be receiving it.
Sub-categories: adult content and erotica; direct erotic chat; non-consensual nudity; intimate-image based abuse; monetised sexual content; indecent exposure; sexualisation.
Child sexual exploitation
Content that sexualises children or enables their abuse. The one category in this domain where any exposure at all is unacceptable rather than a matter of degree.
Sub-categories: child sexual abuse material; sexualisation of children; grooming and inappropriate adult-child relationships; child endangerment; non-sexual child abuse.
Self-harm and mental health content
Output that encourages, instructs or normalises harm to oneself, including responses that affirm a user's destructive intentions rather than challenging them.
Sub-categories: suicide; non-suicidal self-injury; eating disorders; dangerous challenges and hoaxes; affirming destructive thoughts and actions.
Illegal and dangerous activity content
Content that enables, encourages or endorses criminal or hazardous acts, including confident advice in areas where being wrong causes real damage.
Content that enables, encourages or endorses criminal or hazardous acts, including confident advice in areas where being wrong causes real damage.
Sub-categories: crimes and illegal activities; illegal and heavily regulated substances; illegal services and exploitation; dangerous use; harmful advice in high-stakes domains such as health, safety, legal and financial.
Inflammatory political and culturally sensitive content
Material that inflames political division or breaches norms that vary by place, so a model judged safe in one market is unsafe in another.
Sub-categories: sensitive politics; subversive or aggressive political opinions; extremist views on political topics; cultural insensitivity; the absence of any universal standard, since what counts as sensitive shifts by culture, region and language.
3. Unequal performance across groups
Lower performance for some languages
Models are trained in a handful of languages and work measurably less well in the rest, mostly because nobody ever built the labelled training data.
Models are trained in a handful of languages and work measurably less well in the rest, mostly because nobody ever built the labelled training data.
Sub-categories: low-resource languages with no systematic labelled datasets, Javanese and its 80 million-plus speakers being the standing example; deprioritisation of languages whose speakers can fall back on English; safety benchmarking conducted almost entirely in English; multilingual erosion.
Lower performance for some social groups
Accuracy that tracks who the user is rather than what they asked.
Accuracy that tracks who the user is rather than what they asked.
Sub-categories: disparities by race, gender, culture, age and disability; dialect and accent handling; speech recognition racial disparities; social class and education background, which are not typically protected characteristics under anti-discrimination law and so leave no route to complaint.
Quality-of-service harms
The system works, just not as well for you.
The system works, just not as well for you.
Sub-categories: alienation, the self-estrangement felt at the point of use; increased labour, the extra time and effort to get the same result; service or benefit loss.
Unfair capability distribution
Performing worse for the group that was already worse off.
Performing worse for the group that was already worse off.
Sub-categories: disparate performance across tasks such as question answering and fact-checking; performance gaps that compound existing disadvantage; public service delivery deterioration.
No opt-out and forced adaptation
Being unable to avoid the system increases both how often and how badly you experience its failures.
Being unable to avoid the system increases both how often and how badly you experience its failures.
Sub-categories: automated agents replacing humans as the only interface to essential services; users modifying how they speak in order to be understood; linguistic barriers preventing access to welfare and social services; inability to choose a human alternative.
Rosie Sherry
CEO & Founder at Ministry of Testing
She/Her
I've been working in the software testing and quality engineering space since the year 2000 whilst also combining it with my love for education and community. It turns out quality, community and education go nicely hand in hand.
🎓 MoT-STEC qualified
Open To
Write
Teach
Speak
Mentor
CV Reviews
Podcasting
Meet at MoTaCon 2026
Sign in
to comment
Manage your entire QA lifecycle in one place. Sync Jira, automate scripts, and use AI to accelerate your testing.
Explore MoT
What I learned about influence by becoming a stakeholder
Boost your career in software testing with the MoT Software Testing Essentials Certificate. Learn essential skills, from basic testing techniques to advanced risk analysis, crafted by industry experts.
Debrief the week in Quality via a community radio show hosted by Simon Tomes and members of the community