AI Governance Library

Red Teaming Artificial Intelligence for Social Good: The PLAYBOOK

A UNESCO playbook, written with Humane Intelligence, that guides organizations and communities through designing and running red teaming exercises on generative AI models to expose bias, stereotypes and technology-facilitated gender-based violence.
Cover of Red Teaming Artificial Intelligence for Social Good: The PLAYBOOK

⚡ Quick Summary

Published by UNESCO (United Nations Educational, Scientific and Cultural Organization), this playbook was written by Dr Rumman Chowdhury, Theodora Skeadas, Dhanya Lakshmi and Sarah Amos of the nonprofit Humane Intelligence, and builds on a red teaming exercise UNESCO ran with senior diplomats and UNESCO staff on the International Day for the Elimination of Violence against Women.

Its purpose is to make AI testing widely accessible: a step-by-step guide for organizations and communities to design and run their own red teaming exercises on generative AI models, with no additional IT skills required. Red teaming is defined as hands-on testing of Gen AI models for flaws and vulnerabilities, using carefully designed prompts in a safe and controlled environment. The playbook separates unintended consequences (embedded bias, stereotypes, the AI bias reinforcement cycle) from intended malicious attacks (non-consensual deepfake content, prompt injection against trust and safety guardrails).

It supplies worked challenges on gender bias in STEM assessment and on violence against women journalists, guidance on validating and analysing results, reporting, follow-up, and a sample red teaming report template, closing with a call to action.

🧩 What’s Covered

The playbook follows the arc of running a red teaming event, from set-up to follow-up.

  • Introduction (pp. 6–9): Gen AI's promise for gender equality set against harms to women and girls, including the estimate that some girls experience their first technology-facilitated gender-based violence (TFGBV) at nine years old, and 58% of young women and girls globally having experienced online harassment; human oversight is cited from UNESCO's 2021 Recommendation on the Ethics of AI.
  • What red teaming is and who it is for (pp. 8–9): a definition of red teaming as intentionally testing Gen AI models to expose vulnerabilities, plus target users: technology and AI practitioners, researchers and academics, government and policy experts, civil society and nonprofits, educators and students, artists and cultural sector professionals, and citizen scientists.
  • Unintended consequences versus intended malicious attacks (pp. 10–11): the AI bias reinforcement cycle diagram and the pathways of harm, with figures including 30% of AI professionals being women and 96% of deepfake videos being non-consensual intimate content.
  • Preparing for red teaming exercises (pp. 12–16): roles in the co-ordination group (senior leadership, subject matter experts, technical experts and evaluators, facilitator and support crew, at about one support person per 20 participants), expert versus public red teaming, in-person, online and hybrid formats, third-party collaboration and psychological safety.
  • Defining challenges and prompts (pp. 17–21): how to narrow a theme, for example "Does AI perpetuate negative stereotypes about scholastic achievement?", with two worked challenges — a fill-in-the-blank maths aptitude prompt comparing Chineme and David, and a prompt requesting insults about a woman journalist that illustrates prompt injection and automation.
  • Turning findings into action (pp. 22–23): validation and analysis tips, analytical tools (Pysentimiento for hate detection, Microsoft Excel for basic analysis, with the DEFCON 2023 event analysing 164,208 messages across 17,469 conversations), reporting and communication, and follow-up after six months or a year.
  • Common challenges and how to overcome them (pp. 24–25): lack of familiarity with red teaming and AI tools, unclear goals, concerns about time and resources, and resistance, with suggested responses.
  • Glossary, author biographies and sample report template (pp. 28–32): definitions from adversarial testing to TFGBV, the four authors' biographies, a bibliography, and a template with objective, methodology, findings, analysis, recommendations and conclusion sections.

💡 Why it matters?

Organizations outside the large AI labs have few routes into evaluating the generative models they use, and this playbook lowers that barrier. It gives civil society, public institutions and community groups a reproducible method for producing evidence about bias and gender-based harms, and for converting that evidence into recommendations for model owners and decision-makers. It positions red teaming findings as inputs to policy, standards, AI audits and ethics reviews, and includes follow-up to check whether model owners incorporated the findings. Guidance on participant well-being, evaluator independence and validation of flagged content makes the method usable where testing touches distressing material.

❓ What’s Missing

The playbook is deliberately non-technical and does not specify platform requirements, budgets, staffing costs or an exercise timeline beyond the ratio of one support person per 20 participants. It recommends consulting a prompt library but supplies only two worked prompts, and both illustrative challenges concern gender bias in STEM assessment and violence against women journalists, so other harm domains are not demonstrated. Red teaming is not mapped to legal obligations or named standards, no thresholds define an unacceptable model response, and the sample report is a blank template rather than a completed example.

👥 Best For

Best for civil society and gender equality organizations, public institutions and educators planning a first red teaming event; subject matter experts and facilitators who need role descriptions, format choices and worked prompts; and policy or research teams gathering evidence on generative AI harms for advocacy, AI audits or ethics reviews.

📄 Source Details

Red Teaming Artificial Intelligence for Social Good: The PLAYBOOK, published by UNESCO (Paris) in 2025, written by Dr Rumman Chowdhury, Theodora Skeadas, Dhanya Lakshmi and Sarah Amos (Humane Intelligence) and edited by Danielle Cliche and Sinéad Andrews. ISBN 978-92-3-100758-3; Open Access under CC-BY-SA 3.0 IGO. The supplied PDF runs to 36 pages including covers, with internal pagination to 32 and the sample report template on internal page 32. The analysis draws on a text extraction of all 36 pages; no URL for the document itself is printed.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.