Publications
For the most up-to-date list of my papers, please see my Google Scholar profile.
* equal contribution
Spotlight
Oral
Runner-up
Best paper
Select High-Impact Papers
- Biderman, Schoelkopf, et al. “Pythia: A suite for analyzing large language models across training and scaling.” International Conference on Machine Learning (ICML). 2023. Oral (top 2.4%). [Paper] [Code] [Models]Impact: Pythia changed how researchers study large language models. It provided a controlled suite of models at multiple scales, all trained on the same public data in a reproducible and transparent way, enabling a wide range of previously impossible experiments. Most importantly, we released 154 checkpoints for every Pythia model—more checkpoints than the rest of the world had previously released for all models of comparable or greater size combined. These complete training trajectories made it possible to study when and how capabilities, representations, memorization, and biases develop rather than examining only finished models. Thousands of papers on interpretability, training dynamics, memorization, and related topics have relied on Pythia as infrastructure for rigorous research.
- Gao, Biderman, Black, Golding, Hoppe, Foster, Phang, He, Thite, Nabeshima, Presser, and Leahy. “The Pile: An 800GB Dataset of Diverse Text for Language Modeling.” arXiv preprint arXiv:2101.00027. 2020. [Paper] [Datasheet] [Website] [Model]Impact: The Pile was the first major openly released dataset designed specifically for training large language models. Its mixture of 22 diverse sources---including scientific papers, code, books, web text, and domain-specific archives---made it possible to study how data composition affects model capabilities in a way that closed training corpora could not support. Many of its design choices, including deliberate inclusion of high-quality technical and scientific sources, became standard in later pretraining datasets. The Pile also enabled an ecosystem of open models and reproducible experiments, making training data itself a visible and inspectable component of language-model research.
- Black*, Biderman*, Hallahan*, et al. “GPT-NeoX-20B: An Open-Source Autoregressive Language Model.” Workshop on Challenges & Perspectives in Creating Large Language Models at ACL. 2022. [Paper] [Code]Impact: GPT-NeoX-20B was, at the time of its release, the largest and most powerful open-source language model in the world. Just as importantly, the project released the GPT-NeoX training framework, giving researchers a practical system for distributed pretraining at the tens-of-billions-of-parameters scale. GPT-NeoX has since become one of the world's most widely used large-model pretraining libraries, supporting numerous open models and making large-scale training accessible beyond a small number of major technology companies. The model's tokenizer was also unusually innovative and became the de facto standard for open language models for several years. Together, these contributions helped establish open, reproducible large-scale language-model training as a durable engineering and research practice.
- O'Brien, Casper, Anthony, Korbak, Kirk, Davies, Mishra, Irving, Gal, and Biderman. “Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs.” International Conference on Learning Representations (ICLR). 2026. Best paper runner-up at the NeurIPS 2025 Workshop on Biosecurity Safeguards for Generative AI. [Paper] [Project Page] [Artifacts]Impact: Deep Ignorance introduced a new approach to open-weight AI safety: preventing a model from acquiring a dangerous capability during pretraining, rather than trying to suppress that capability afterward. The work showed that carefully filtering training data can selectively impair learning in a targeted domain while preserving broad model usefulness. This is especially important for open-weight systems, where safeguards applied only at inference time can be removed by anyone with access to the weights. The paper helped establish data curation as a serious safety intervention and opened a research direction focused on making safety properties durable under adversarial fine-tuning and unrestricted deployment.
- Biderman, Khan, Mireshghallah, Arnett, Barez, and Saphra. “Position: Don't Just “Fix it in Post”: A Science of AI Must Study Training Dynamics.” International Conference on Machine Learning (ICML). 2026. Oral (top 0.7%). [Paper] [OpenReview]Impact: Most research evaluates models only after training is complete, obscuring the process that produced their capabilities, failures, and biases. This position paper argues that AI should instead be studied as a developmental science: researchers should observe models throughout training, identify when important behaviors emerge, and learn how data and optimization shape their trajectories. That shift would make it possible to move beyond diagnosing finished systems toward predicting and intervening on model behavior while it is still forming. The paper lays out a research program for treating training dynamics as a central object of study rather than an implementation detail.
- Kandpal*, Lester*, Raffel*, Majstorovic, Biderman, et al. “The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text.” Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track. 2025. [Paper] [Artifacts]Impact: The Common Pile demonstrated that a genuinely large, high-quality language-model corpus can be built entirely from public-domain and openly licensed text. This matters because the field had increasingly treated indiscriminate scraping of copyrighted material as a practical prerequisite for competitive model training. By assembling eight terabytes of legally reusable data and documenting its provenance, composition, and processing, the project established a viable foundation for models that can be studied, reproduced, and redistributed more freely. It also created critical infrastructure for an open data commons at a moment when access to training data was rapidly becoming more restricted.
- Peng, Alcaide, Anthony, et al. (incl. Biderman). “RWKV: Reinventing RNNs for the Transformer Era.” Findings of the Association for Computational Linguistics (EMNLP). 2023. [Paper] [Blog] [Model] [Code]Impact: RWKV was the first non-transformer architecture scaled beyond 7 billion parameters to match the performance of similarly sized transformers. This result provided the first compelling evidence at modern language-model scale that transformers were not the only viable architecture for high-performing large language models. RWKV combines parallelizable training with recurrent inference, enabling generation with constant memory per token rather than an ever-growing attention cache. Beyond the specific architecture, the work reopened serious investigation into alternative sequence models and helped motivate the broader wave of efficient, recurrent, and long-context architectures that followed.
- Ahdritz, et al. (incl. Biderman). “OpenFold: Retraining AlphaFold2 yields new insights into its learning mechanisms and capacity for generalization.” Nature Methods, 21(8). 2024. [Paper]Impact: AlphaFold2 represented a major advance in protein-structure prediction, but its initial closed release sharply limited its value as a platform for scientific research. To make this technology genuinely available to the scientific community, we partnered with academic researchers to train an open-source replication. Shortly before OpenFold was released—and in response to the project—Google DeepMind released AlphaFold2 under a non-commercial license. OpenFold went substantially further: we released the model under an open-source license together with its training code and partially trained checkpoints. This made it possible not only to use and modify the system, but also to study how its capabilities and generalization emerge throughout training, turning a landmark result into an inspectable foundation for further research.
- Biderman, Prashanth, Sutawika, Schoelkopf, Anthony, Purohit, and Raff. “Emergent and predictable memorization in large language models.” Neural Information Processing Systems (NeurIPS). 2023. [Paper]Impact: This paper changed how memorization in language models can be studied by tracking individual training examples across checkpoints rather than measuring memorization only in a finished model. It showed that memorization is highly structured: examples that appear unmemorized early in training can later undergo sharp transitions, and those outcomes can be predicted from earlier training behavior. This connects privacy and data governance questions to the dynamics of optimization itself. More broadly, the work demonstrated that intermediate checkpoints contain actionable information about a model's eventual behavior, supporting the idea that important properties of final models can be forecast before training is complete.
- Crowson*, Biderman*, Kornis, Stander, Hallahan, Castricato, and Raff. “VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance.” European Conference on Computer Vision (ECCV). 2022. [Paper] [Demo] [Code]Impact: VQGAN-CLIP showed that powerful open-ended text-to-image generation could emerge by composing existing open models rather than training a monolithic system from scratch. By using CLIP to guide images produced through a VQGAN latent space, the work made flexible natural-language-driven generation and editing accessible to artists, researchers, and independent developers. It spread rapidly through the creative-coding community and helped catalyze public interest in modern generative art before diffusion models became dominant. The project is an influential example of how open components, recombined imaginatively, can unlock capabilities beyond those envisioned in their original training.
All Publications
Academic Publications
2026 and currently under review
- Chang*, Arnett*, et al. (incl. Biderman). “Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures.” Under review. 2026. [Paper]
- Batzner*, Nelaturu*, et al. (incl. Biderman). “Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results.” Under review. 2026. [Paper]
- Ghosh et al. (incl. Biderman). “Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting.” Under review. 2026. [Paper] [Project Page]
- Cesista, Crowson, Simal, and Biderman. “LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold.” Under review. 2026. [Paper]
- Barez, Mireshghallah, Osborne, Biderman, Lazar, and Trager. “What AI Governance Needs from Mechanistic Auditing.” Technical AI Governance Workshop at ICML. 2026.
- Quirke, Jaburi, Johnston, Li, Martres, Paulo, Gupta, Biderman, and Belrose. “Bergson: An Open Source Library for Data Attribution.” Under review. 2026. [Paper] [Code]
- Matlin, et al. (incl. Biderman). “Where Does Social Reasoning Come From? Capability Provenance in Language Models.” Conference on Language Modeling (COLM). 2026. [Paper]
- Ammanamanchi, Bhat, and Biderman. “What Helps Agentic Lean Provers? A Trace-Level Attribution Study.” Under review. 2026. [Paper]
- Biderman, Khan, Mireshghallah, Arnett, Barez, and Saphra. “Position: Don't Just “Fix it in Post”: A Science of AI Must Study Training Dynamics.” International Conference on Machine Learning (ICML). 2026. Oral (top 0.7%). [Paper] [OpenReview]
- Ammanamanchi, Bhat, and Biderman. “Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving.” International Conference on Machine Learning (ICML). 2026. [Paper] [Code]
- Akhtar*, Reuel*, et al. (incl. Biderman). “When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation.” International Conference on Machine Learning (ICML). 2026. [Paper]
- Schaeffer, et al. (incl. Biderman). “Quantifying the Effect of Test Set Contamination on Generative Evaluations.” International Conference on Machine Learning (ICML). 2026. [Paper]
- Crawford, Khanna, Lu, Wagoner, Biderman, Nguyen, and Raff. “Adversarial Samples Are Not Created Equal.” Under review. 2026. [Paper]
- Kamachee, Casper, Ding, Yew, Reuel, Biderman, and Hadfield-Menell. “Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns.” Under review. 2026. [Paper]
- Reuel, et al. (incl. Biderman). “Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations.” International Conference on Machine Learning (ICML). 2026. [Paper]
- O'Brien, Casper, Anthony, Korbak, Kirk, Davies, Mishra, Irving, Gal, and Biderman. “Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs.” International Conference on Learning Representations (ICLR). 2026. Best paper runner-up at the NeurIPS 2025 Workshop on Biosecurity Safeguards for Generative AI. [Paper] [Project Page] [Artifacts]
2025
- Kandpal*, Lester*, Raffel*, Majstorovic, Biderman, et al. “The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text.” Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track. 2025. [Paper] [Artifacts]
- Bradshaw, Spangher, Biderman, and Colton. “The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity.” NeurIPS Creative AI Track. 2025. [Paper]
- Arnett, Chang, Biderman, and Bergen. “Explaining and Mitigating Crosslingual Tokenizer Inequities.” Neural Information Processing Systems (NeurIPS). 2025. [Paper]
- Bradshaw, Fan, Spangher, Biderman, and Colton. “Scaling Self-Supervised Representation Learning for Symbolic Piano Performance.” International Society for Music Information Retrieval Conference (ISMIR). 2025. [Paper] [Code]
- Peng, et al. (incl. Biderman). “RWKV-7 ‘Goose' with Expressive Dynamic State Evolution.” Conference on Language Modeling (COLM). 2025. [Paper] [Models] [Code]
- Son, Hong, Fan, Nam, Ko, Lim, Song, Choi, Paulo, Yu, and Biderman. “When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research.” arXiv preprint arXiv:2505.11855. 2025. [Paper] [Data]
- Schaeffer, et al. (incl. Biderman). “Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?” International Conference on Machine Learning (ICML). 2025. [Paper]
- Son, Lee, Kim, Kim, Muennighoff, Choi, Park, Yoo, and Biderman. “KMMLU: Measuring Massive Multitask Language Understanding in Korean.” Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). 2025. [Paper] [Data]
- Longpre*, Biderman*, et al. “The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources.” Transactions of Machine Learning Research (TMLR). 2025. [Paper]
- van der Wal, Lesci, Muller-Eberstein, Saphra, Schoelkopf, Zuidema, and Biderman. “PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs.” International Conference on Learning Representations (ICLR). 2025. [Paper] [Code] [Models]
- Longpre, et al. (incl. Biderman). “Consent in crisis: The rapid decline of the AI data commons.” International Conference on Learning Representations (ICLR). 2025. [Paper]
- Longpre, et al. (incl. Biderman). “Bridging the data provenance gap across text, speech and video.” International Conference on Learning Representations (ICLR). 2025. [Paper]
- Prashanth*, Deng*, O'Brien*, V*, et al. (incl. Biderman). “Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon.” International Conference on Learning Representations (ICLR). 2025. [Paper]
- Sharkey*, Chughtai*, et al. (incl. Biderman). “Open Problems in Mechanistic Interpretability.” Transactions of Machine Learning Research (TMLR). 2025. [Paper]
2024
- Forde*, Zhang*, et al. (incl. Biderman). “Re-Evaluating Evaluation for Multilingual Summarization.” Conference on Empirical Methods in Natural Language Processing (EMNLP). 2024. [Paper]
- Biderman*, Schoelkopf*, Sutawika*, et al. “Lessons from the Trenches on Reproducible Evaluation of Language Models.” arXiv preprint arXiv:2405.14782. 2024. [Paper]
- Alam, Oberle, Raff, Biderman, Oates, and Holt. “A Walsh Hadamard Derived Linear Vector Symbolic Architecture.” Neural Information Processing Systems (NeurIPS). 2024. [Paper] [Code]
- Tigges, Hanna, Yu, and Biderman. “LLM Circuit Analyses Are Consistent Across Training and Scale.” Neural Information Processing Systems (NeurIPS). 2024. [Paper]
- Peng, Goldstein, Anthony, et al. (incl. Biderman). “Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.” Conference on Language Modeling (COLM). 2024. [Paper] [Models] [Code]
- Anthony, Hatef, Narayanan, Biderman, Bekman, Yin, Shafi, Subramoni, and Panda. “The Case for Co-Designing Model Architectures with Hardware.” International Conference on Parallel Processing (ICPP). 2024. [Paper]
- Ahdritz, et al. (incl. Biderman). “OpenFold: Retraining AlphaFold2 yields new insights into its learning mechanisms and capacity for generalization.” Nature Methods, 21(8). 2024. [Paper]
- McMillan-Major*, De Toni*, et al. (incl. Biderman). “Documenting Geographically and Contextually Diverse Language Data Sources.” Northern European Journal of Language Technology, 10:50--77. 2024. [Paper]
- Kapoor, et al. (incl. Biderman). “Position: On the Societal Impact of Open Foundation Models.” International Conference on Machine Learning (ICML). 2024. Oral (top 1.5%). [Paper]
- Strander, Yu, Fan, and Biderman. “Grokking group multiplication with cosets.” International Conference on Machine Learning (ICML). 2024. [Paper]
- Spangher, Sanchez, Fan, Levi, and Biderman. “Stay on Topic with Classifier-Free Guidance.” International Conference on Machine Learning (ICML). 2024. Spotlight (top 3.5%). [Paper]
- Azerbayev, Schoelkopf, Paster, Dos Santos, McAleer, Jiang, Deng, Biderman, and Welleck. “Llemma: An Open Language Model for Mathematics.” International Conference on Learning Representations (ICLR). 2024. [Paper]
- Alam, Raff, Biderman, Oates, and Holt. “Holographic global convolutional networks for long-range prediction tasks in malware detection.” International Conference on Artificial Intelligence and Statistics (AISTATS). 2024. [Paper]
- Zhang, Tigges, Zhang, Biderman, Raginsky, and Ringer. “Transformer-based models are not yet perfect at learning to emulate structural recursion.” Transactions of Machine Learning Research (TMLR). 2024. [Paper]
2023
- Ruis, Khan, Biderman, Hooker, Rocktäschel, and Grefenstette. “The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs.” Neural Information Processing Systems (NeurIPS). 2023. Spotlight (top 3%). [Paper]
- Havrilla, Zhuravinskyi, Phung, Tiwari, Tow, Biderman, Anthony, and Castricato. “trlX: A framework for large scale reinforcement learning from human feedback.” Conference on Empirical Methods in Natural Language Processing (EMNLP). 2023. [Code]
- Peng, Alcaide, Anthony, et al. (incl. Biderman). “RWKV: Reinventing RNNs for the Transformer Era.” Findings of the Association for Computational Linguistics (EMNLP). 2023. [Paper] [Blog] [Model] [Code]
- Belrose, Schneider-Joseph, Ravfogel, Cotterell, Raff, and Biderman. “LEACE: Perfect linear concept erasure in closed form.” Neural Information Processing Systems (NeurIPS). 2023. [Paper]
- Biderman, Prashanth, Sutawika, Schoelkopf, Anthony, Purohit, and Raff. “Emergent and predictable memorization in large language models.” Neural Information Processing Systems (NeurIPS). 2023. [Paper]
- Biderman, Schoelkopf, et al. “Pythia: A suite for analyzing large language models across training and scaling.” International Conference on Machine Learning (ICML). 2023. Oral (top 2.4%). [Paper] [Code] [Models]
- Alam, Raff, Biderman, Oats, and Holt. “Recasting Self-Attention with Holographic Reduced Representations.” International Conference on Machine Learning (ICML). 2023. [Paper] [Code]
- Belrose, et al. (incl. Biderman). “Eliciting latent predictions from transformers with the tuned lens.” arXiv preprint arXiv:2303.08112. 2023. [Paper] [Code]
- Muennighoff, et al. (incl. Biderman*). “Crosslingual Generalization through Multitask Finetuning.” Annual Meeting of the Association for Computational Linguistics (ACL). 2023. Spotlight (top 3%). [Paper] [Models] [Code]
- Yong, et al. (incl. Biderman*). “BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting.” Annual Meeting of the Association for Computational Linguistics (ACL). 2023. [Paper]
- Piktus, et al. (incl. Biderman). “GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration.” Annual Meeting of the Association for Computational Linguistics (ACL). 2023. [Paper] [Code] [Demo]
2022 and earlier
- Laurençon*, et al. (incl. Biderman*). “The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset.” Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track. 2022. Oral. [Paper] [Explorer]
- Fries, et al. (incl. Biderman*). “BigBio: A Framework for Data-Centric Biomedical Natural Language Processing.” Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track. 2022. [Paper] [Code]
- Phang, Bradley, Gao, Castricato, and Biderman. “EleutherAI: Going Beyond Open Science to Science in the Open.” Workshop on Broadening Research Collaborations in ML at NeurIPS. 2022. [Paper]
- Sanh*, Webson*, Raffel*, Bach*, et al. (incl. Biderman). “Multitask Prompted Training Enables Zero-Shot Task Generalization.” International Conference on Learning Representations (ICLR). 2022. [Paper] [Model]
- Le Scao*, et al. (incl. Biderman*). “Bloom: A 176b-parameter open-access multilingual language model.” Transactions of Machine Learning Research (TMLR). 2022. [Paper] [Model]
- Srivastava, et al. (incl. Biderman). “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.” Transactions of Machine Learning Research (TMLR). 2022. Outstanding Paper Finalist (top 5 out of 956). [Paper] [Code]
- Biderman and Raff. “Fooling MOSS detection with pretrained language models.” ACM International Conference on Information and Knowledge Management (CIKM). 2022. [Paper]
- Crowson*, Biderman*, Kornis, Stander, Hallahan, Castricato, and Raff. “VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance.” European Conference on Computer Vision (ECCV). 2022. [Paper] [Demo] [Code]
- Jernite*, Nguyen*, Biderman, et al. “Data Governance in the Age of Large-Scale Data-Driven Language Technology.” ACM Conference on Fairness, Accountability, and Transparency (FAccT). 2022. [Paper]
- Black*, Biderman*, Hallahan*, et al. “GPT-NeoX-20B: An Open-Source Autoregressive Language Model.” Workshop on Challenges & Perspectives in Creating Large Language Models at ACL. 2022. [Paper] [Code]
- Hesslow*, Le Scao*, Saulnier*, et al. (incl. Biderman). “What Language Model to Train if You Have One Million GPU Hours?” Findings of the Association for Computational Linguistics (EMNLP). 2022. [Paper]
- Talat, Névéol, Biderman, et al. “You Reap What You Sow: On the Challenges of Bias Evaluation under Multilingual Settings.” Workshop on Challenges & Perspectives in Creating Large Language Models at ACL. 2022. [Paper]
- Caswell*, Kreutz*, et al. (incl. Biderman). “Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets.” Transactions of the Association for Computational Linguistics (TACL). 2022. [Paper]
- Biderman, Bicheno, and Gao. “Datasheet for the Pile.” arXiv preprint arXiv:2201.07311. 2022. [Paper]
- Alcaide, Biderman, Telenti, and Maher. “Massively Parallel Natural Extension of Reference Frame for Efficient Internal to Cartesian Conversion.” Journal of Computational Chemistry. 2021. [Paper] [Code]
- Matiana*, Smith*, Teehan*, Castricato*, Biderman*, Gao, and Fraizer. “Cut the CARP: Fishing for zero-shot story evaluation.” arXiv preprint arXiv:2110.03111. 2021. [Paper]
- Castricato*, Biderman*, Cardona-Rivera, and Thue. “Towards a Model-theoretic View of Narratives.” Workshop on Narrative Understanding at NAACL-HLT. 2021. [Paper]
- Gao, Biderman, Black, Golding, Hoppe, Foster, Phang, He, Thite, Nabeshima, Presser, and Leahy. “The Pile: An 800GB Dataset of Diverse Text for Language Modeling.” arXiv preprint arXiv:2101.00027. 2020. [Paper] [Datasheet] [Website] [Model]
- Biderman and Scheirer. “Pitfalls in Machine Learning Research: Reexamining the Development Cycle.” “I Can't Believe It's Not Better!” Workshop at NeurIPS. 2020. [Paper]
- Churchill, Biderman, and Herrick. “Magic: The Gathering is Turing Complete.” International Conference on Fun with Algorithms (FUN). 2020. [arXiv] [Demo]
- Biderman. “Neural Networks on Groups.” arXiv preprint arXiv:1907.03742. 2020. [Paper]
- Biderman. “Magic: the Gathering is as Hard as Arithmetic.” arXiv preprint arXiv:2003.05119. 2019. [Paper]
Books
- Raff, Farris, and Biderman. How Large Language Models Work. Simon and Schuster, 2025. [Book]
Non-Archival Workshops and Conferences
- O'Brien, Casper, Anthony, Korbak, Kirk, Davies, Mishra, Irving, Gal, and Biderman. “Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs.” Workshop on Biosecurity Safeguards for Generative AI at NeurIPS 2025. 2026. Best paper runner-up. [Paper] [Project Page] [Artifacts]
- Biderman, Mickel, and Abbasi. “Write Code that People Want to Use.” Workshop on Championing Open-source DEvelopment in ML (CODEML) at ICML. 2025. [Paper]
- Baack, Biderman, Odrozek, et al. “Towards Best Practices for Open Datasets for LLM Training.” Free and Open Source Software Developers' European Meeting (FOSDEM). 2025. [Paper]
- O'Mahony, Grinsztajn, Schoelkopf, and Biderman. “Attributing mode collapse in the fine-tuning of large language models.” Workshop on Mathematical and Empirical Understanding of Foundation Models at ICLR. 2024. [Paper]
- Hesslow*, Le Scao*, Saulnier*, et al. (incl. Biderman). “What Language Model to Train if You Have One Million GPU Hours?” Workshop on Challenges & Perspectives in Creating Large Language Models at ACL. 2022. [Paper]
- Caswell*, Kreutz*, et al. (incl. Biderman). “Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets.” Workshop on African Natural Language Processing at EACL. 2021. [Paper]
- Masad, Biderman, and Shishkoff. “Predicting Crisis Behavior with Reinforcement Learning.” Military Operations Research Society's Emerging Techniques Forum. 2019. Eugene P. Visco Prize for best research by a junior analyst. [Slides]
- Masad, Biderman, Shishkoff, and Baird. “Reinforcement Learning in Conflict Escalation Games.” Annual Meeting of the Society for Political Methodology (PolMeth). 2018. [Abstract] [Poster]
- Biderman, Masad, and Lawson. “How to be Wrong, but Useful: A Case Study on Tool Selection in Social Network Analysis.” North American Social Networks Conference. 2018. [Slides]