Debiasing Algorithm through Model Adaptation

PID

Debiasing Algorithm through Model Adaptation (DAMA) is based on guarding stereotypical gender signals and model editing. DAMA is performed on specific modules prone to convey gender bias, as shown by causal tracing. Our novel method effectively reduces gender bias in LLaMA models in three diagnostic tests: generation, coreference (WinoBias), and stereotypical sentence likelihood (StereoSet). The method does not change the model’s architecture, parameter count, or inference cost. We have also shown that the model’s performance in language modeling and a diverse set of downstream tasks is almost unaffected. This package contains both the source codes and English, English-to-Czech, and English-to-German datasets.

Identifier
PID http://hdl.handle.net/11234/1-5847
Related Identifier https://openreview.net/pdf?id=XIZEFyVGC9
Related Identifier https://ufal.mff.cuni.cz/grants/genderbias
Metadata Access http://lindat.mff.cuni.cz/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:lindat.mff.cuni.cz:11234/1-5847
Provenance
Creator Limisiewicz, Tomasz; Mareček, David; Musil, Tomáš
Publisher Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Publication Year 2025
Rights The MIT License (MIT); http://opensource.org/licenses/mit-license.php; PUB
OpenAccess true
Contact lindat-help(at)ufal.mff.cuni.cz
Representation
Language English; Czech; German
Resource Type toolService
Format application/zip; application/octet-stream; downloadable_files_count: 1
Discipline Linguistics