Study results: Detecting gendered harm with Meedan’s ALTO dataset

Our new research proves that if classifiers are trained by communities that are encountering TFGBV first-hand, they will improve.

Read the Executive Summary of our research on detecting TFGBV in the Larger World.
No items found.

From sexual harassment to doxxing to other coordinated forms of abuse, technology-facilitated gender-based violence (TFGBV) has been well documented across major social media platforms over the past decade. Yet the technological interventions designed to tackle this problem have focused overwhelmingly on English and other European-language content in North America and Europe, leaving most of the world with far weaker online protections.

This inequity has driven Meedan’s program and research teams to study the unique linguistic and contextual dynamics of TFGBV in the Larger World. Since 2025, we have tracked and analyzed TFGBV in African French, Levantine Arabic, and Swahili on the Internet across various platforms supported by regional community organizations. This mixed-methods research project was intended to deepen public understanding of how TFGBV manifests in major languages of Western Asia and Africa. The research also helped improve machine learning solutions to reduce the spread of TFGBV and led us to develop ALTO: African and Levantine Tech-facilitated Gender-based Violence (TFGBV) Corpus. 

How do we define TFGBV?

TFGBV is a prevalent and growing form of harm that targets women, girls, and gender-diverse people on the basis of their gender and sexual identity. While the official UN definition we adopted describes TFGBV as gendered content, harm, and the use of technology, we found an incredibly nuanced range of manifestations that required careful deliberation with annotators. The project also set out to understand the machine learning resources currently available to detect TFGBV beyond the languages of the Global North in Levantine Arabic, African French, and Swahili – and we found surprisingly few. In this blog post, we want to share some findings from this work and discuss where to take it next. 

TFGBV Resource Survey

We conducted a thorough survey of the datasets and models available for TFGBV detection. Annotated datasets are central to training new models and for benchmarking (i.e. evaluating) past models.

We reviewed 140 harm-detection datasets from academic sources and the open-source ML repository HuggingFace. After removing datasets that did not meet the language, gendered harm, or accessibility requirements for our study, we were left with only four datasets. Interestingly, all of these datasets were multilingual: all four supported Arabic (mixed dialects), three supported French (no explicit dialect), and only one supported Swahili, containing only 17 positive examples of TFGBV. The scarcity of resources across languages and within datasets highlights the importance of this work and demonstrates the need for further research in this area.

Harm Detection Models

The models used for TFGBV and broader harmful text detection fall into two categories: encoder-based classifiers trained on annotated data, and safety-finetuned decoder-only models (e.g., GPT-OSS-Safeguard), which are LLMs trained to classify content based on open-ended content policies. 

We selected several established models from leading platforms, including:

  • GPT-OSS-Safeguard 20B and 120B from OpenAI [1]
  • Learning From The Worst (LFTW) from Bertie Vidgen and Facebook [2]
  • Google Perspective API. [3]

We assessed performance across language settings using the new ALTO dataset and the surveyed past datasets. 

Experimental Results

We ran experiments on the newly established ALTO dataset and the surveyed datasets to understand the current model’s performance on TFGBV across our three language settings and observe how fine-tuning on ALTO can improve performance. We introduced the active learning method for building ALTO in an earlier blog post.

ALTO - TFGBV Bench - TFGBV
LangModel PrecisionRecallF1 PrecisionRecallF1
arOSS-SG-120B0.920.670.780.370.420.39
OSS-SG-20B0.830.640.720.360.540.43
LFTW0.690.730.710.290.430.34
P-API0.760.710.730.130.390.19
MARBERTv20.830.770.800.500.580.54
AfroXLMR-L-76L (M)0.780.730.750.490.810.61
frOSS-SG-120B0.630.700.660.200.580.30
OSS-SG-20B0.650.760.700.170.650.26
LFTW0.630.540.580.060.650.11
P-API0.600.740.660.040.770.08
CamemBERT-large0.840.830.840.330.500.40
AfroXLMR-L-76L (M)0.810.760.790.230.420.30
sw*OSS-SG-120B0.720.480.580.000.000.00
OSS-SG-20B0.670.540.600.000.000.00
LFTW0.570.420.480.000.000.00
AfroXLMR-base0.760.580.660.110.330.17
AfroXLMR-L-76L (M)0.660.750.700.100.330.15

Swipe sideways to see all scores →

Table 1 - The classification performance of TFGBV by model on the ALTO and surveyed (bench) datasets for Levantine Arabic (ar), African French (fr) and Swahili (sw). The bolded numbers represent the best metric per model for each language. The shaded rows represent the models fine-tuned on ALTO, and (M) denotes multilingual training. *Swahili bench dataset contained only 3 positive TFGBV examples which contributed to its zero performance.

We report our experimental results on TFGBV detection using three metrics: precision, recall, and F1. Precision tells us how often the model is correct within all of its positive TFGBV predictions. Recall tells us the percentage of TFGBV predictions caught within the entire dataset. The F1 takes a balance of the two and is the number we focus on to strike a balanced classifier.

Several interesting results appear from our experimental results shown in Table 1. Overall, we see that fine-tuning on the ALTO dataset improves the TFGBV detection performance on nearly all metrics beyond current state-of-the-art models. The improvement holds across both the ALTO dataset and the surveyed dataset, showing that training on our sample improves performance on out-of-distribution past datasets. 

The result also showed that existing general harm detection models we tested struggle with TFGBV across both the ALTO and surveyed datasets, with a noticeable pattern across the languages. Overall, Arabic saw the highest F1 performance by current models, while French was second and Swahili saw the lowest scores across both datasets. The OpenAI OSS-Safeguard model was the best off-the-shelf model across the three languages (ar-F1=0.78, fr-F1=0.70, sw-F1=0.60), but it trades overall recall for higher precision. The best-performing models for each language, as measured by these metrics, are the ALTO fine-tuned models.

Conclusion

The ALTO dataset represents some of the first TFGBV resources developed for Levantine Arabic, African French, and Swahili. Our experiments demonstrated that existing general harm detection models struggle to accurately identify TFGBV across these languages. Fine-tuning models on the new ALTO corpus substantially improved detection performance. The findings from this work further underscore the critical need for continued targeted, localized resources to combat online gendered harm effectively.

-

Footnotes

[1] https://openai.com/index/introducing-gpt-oss-safeguard/‍

[2] https://huggingface.co/facebook/roberta-hate-speech-dynabench-r4-target, English only but given its established nature in hate speech research we included it to operate on the English translations

[3] https://www.perspectiveapi.com/, only available for French and Arabic without English translation for comparability. 

-

This work was made possible by support from the Christchurch Call Foundation as a part of the Project Catalyst consortium and the United Nations Population Fund (UNFPA), whose 2021 definition of TFGBV provided the baseline for our approach. The views expressed are those of the authors and do not necessarily reflect the views of UNFPA, the United Nations or any of its affiliated organizations.

Highlighted numbers

No items found.

Funded by

No items found.

Partners

No items found.