publications
"*" denotes co-first authorship
2026
- ACL Findings
When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype MitigationYahan Zheng, John J. Guerrerio, Soroush Vosoughi, and Weicheng MaIn Findings of the Association for Computational Linguistics: ACL 2026, 2026Preprocessing-based methods for stereotype mitigation, such as pre-/post-training on debiased corpora, are widely used in NLP. While these approaches reduce measurable stereotypes for targeted groups, we find they often induce unintended shifts-side effects, where stereotyping or counter-stereotyping can increase relative to neutral baselines for other demographics, including across unrelated demographic categories. We demonstrate these side effects across two model families (encoder-only and decoder-only), multiple preprocessing strategies (removing stereotypical sentences, removing group mentions, and swapping group references), and both pre- and post-training at different data scales on Wikipedia. Standard benchmarks frequently miss these shifts. Using attention-rollout analysis, we observe that such side effects are not accompanied by large changes in attention flow, complicating mechanistic explanations. We discuss implications for evaluation, provide actionable diagnostics, and argue for side-effect-aware, transparent mitigation practices.
@inproceedings{whenDebiasingBackfires, title = {When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation}, author = {Zheng, Yahan and Guerrerio, John J. and Vosoughi, Soroush and Ma, Weicheng}, booktitle = {Findings of the Association for Computational Linguistics: ACL 2026}, year = {2026}, }
2025
- EMNLP
Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM CollaborationWeicheng Ma*, John J. Guerrerio*, and Soroush VosoughiIn Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025Research on stereotypes in large language models (LLMs) has largely focused on English-speaking contexts, due to the lack of datasets in other languages and the high cost of manual annotation in underrepresented cultures. To address this gap, we introduce a cost-efficient human-LLM collaborative annotation framework and apply it to construct EspanStereo, a Spanish-language stereotype dataset spanning multiple Spanish-speaking countries across Europe and Latin America. EspanStereo captures both well-documented stereotypes from prior literature and culturally specific biases absent from English-centric resources. Using LLMs to generate candidate stereotypes and in-culture annotators to validate them, we demonstrate the framework’s effectiveness in identifying nuanced, region-specific biases. Our evaluation of Spanish-supporting LLMs using EspanStereo reveals significant variation in stereotypical behavior across countries, highlighting the need for more culturally grounded assessments. Beyond Spanish, our framework is adaptable to other languages and regions, offering a scalable path toward multilingual stereotype benchmarks. This work broadens the scope of stereotype analysis in LLMs and lays the groundwork for comprehensive cross-cultural bias evaluation.
@inproceedings{espanStereo, title = {Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration}, author = {Ma, Weicheng and Guerrerio, John J. and Vosoughi, Soroush}, booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing}, year = {2025}, }
2024
- NeurIPS BDU
Decision-Driven Calibration for Cost-Sensitive Uncertainty QuantificationGreg Canal, Vladimir Leung, John J. Guerrerio, Philip Sage, and I-Jeng WangIn NeurIPS 2024 Workshop on Bayesian Decision-making and Uncertainty, 2024In recent years, the ability of artificial intelligence (AI) systems to quantify their uncertainty has become paramount in building trustworthy AI. In standard uncertainty quantification (UQ), AI uncertainty is calibrated such that the confidence of its predictions matches the statistics of the underlying data distribution. However, this method of calibration does not take into consideration the direct influence of UQ on the subsequent actions taken by downstream decision-makers. Here we demonstrate an alternate, decision-driven method of UQ calibration that explicitly minimizes the incurred costs of downstream decisions. After formulating decision-driven calibration as an optimization problem with respect to a known decision-maker, we show in a simulated search-and-rescue scenario how decision-driven temperature scaling can lead to lower incurred decision costs.
@inproceedings{decisionDrivenCalibration, title = {Decision-Driven Calibration for Cost-Sensitive Uncertainty Quantification}, author = {Canal, Greg and Leung, Vladimir and Guerrerio, John J. and Sage, Philip and Wang, I-Jeng}, booktitle = {NeurIPS 2024 Workshop on Bayesian Decision-making and Uncertainty}, year = {2024}, }
2023
- IT Pro
Labeling Software Security VulnerabilitiesIrena Bojanova and John J. GuerrerioIT Professional, 2023Labeling software security vulnerabilities would benefit greatly modern artificial intelligence cybersecurity research. The National Vulnerability Database (NVD) partially achieves this via assignment of Common Weakness Enumeration (CWE) entries to Common Vulnerabilities and Exposures (CVE) entries. In this work, we explore utilization of the Bugs Framework (BF) formalism for systematic and comprehensive CVE labeling. We specify all memory related CWEs via BF weaknesses and examine the suitability of these formalisms to describe the corresponding CVEs mapped by NVD. We also identify the similarities and overlaps in CWEs that introduce ambiguities in NVD assignments.
@article{bugsFramework, title = {Labeling Software Security Vulnerabilities}, author = {Bojanova, Irena and Guerrerio, John J.}, journal = {IT Professional}, volume = {25}, number = {5}, pages = {64--70}, year = {2023}, doi = {10.1109/MITP.2023.3314368}, } - NAR
LitCovid in 2022: an information resource for the COVID-19 literatureQingyu Chen*, Alexis Allot*, Robert Leaman, Chih-Hsuan Wei, Elaheh Aghaarabi, John J. Guerrerio, Lilly Xu, and Zhiyong LuNucleic Acids Research, 2023LitCovid (https://www.ncbi.nlm.nih.gov/research/coronavirus/)—first launched in February 2020—is a first-of-its-kind literature hub for tracking up-to-date published research on COVID-19. The number of articles in LitCovid has increased from 55 000 to ∼300 000 over the past 2.5 years, with a consistent growth rate of ∼10 000 articles per month. In addition to the rapid literature growth, the COVID-19 pandemic has evolved dramatically. For instance, the Omicron variant has now accounted for over 98% of new infections in the United States. In response to the continuing evolution of the COVID-19 pandemic, this article describes significant updates to LitCovid over the last 2 years. First, we introduced the long Covid collection consisting of the articles on COVID-19 survivors experiencing ongoing multisystemic symptoms, including respiratory issues, cardiovascular disease, cognitive impairment, and profound fatigue. Second, we provided new annotations on the latest COVID-19 strains and vaccines mentioned in the literature. Third, we improved several existing features with more accurate machine learning algorithms for annotating topics and classifying articles relevant to COVID-19. LitCovid has been widely used with millions of accesses by users worldwide on various information needs and continues to play a critical role in collecting, curating and standardizing the latest knowledge on the COVID-19 literature.
@article{litCovid, title = {{LitCovid} in 2022: an information resource for the {COVID-19} literature}, author = {Chen, Qingyu and Allot, Alexis and Leaman, Robert and Wei, Chih-Hsuan and Aghaarabi, Elaheh and Guerrerio, John J. and Xu, Lilly and Lu, Zhiyong}, journal = {Nucleic Acids Research}, volume = {51}, number = {D1}, pages = {D1512--D1518}, year = {2023}, doi = {10.1093/nar/gkac1005}, }