Shipping AI Oct 2025 · #72 most cited · 76 citations

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

Alexandra Souly et al.

arXiv:2510.07192

In short

In the largest pre-training poisoning experiments to date, about 250 malicious documents were enough to plant a backdoor in models from 600M to 13B parameters, regardless of how much clean data they trained on. The same held during fine-tuning.

Why it matters

Bigger training sets do not dilute poisoning, so attacks may be easier than assumed.

Read first

The 4 Field Guide ideas this paper leans on.

Starting from scratch? The full route 9 ideas · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training Data · read first ✓ understood

    The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.

  3. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  4. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  5. Unsupervised Learning ✓ understood

    Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.

  6. Self-Supervised Learning ✓ understood

    Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).

  7. Pre-training · read first ✓ understood

    Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.

  8. Data Poisoning · read first ✓ understood

    Corrupting training data to manipulate model behavior or introduce vulnerabilities.

  9. Backdoor Attack · read first ✓ understood

    Maliciously training models to behave normally except when specific triggers are present.

In the frontier

Rank
#72 of 100
Citations
76
as of Aug 9, 2026
Published
Oct 2025

Topics: Safety and alignment , Data and synthetic generation , Interpretability and analysis

Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026

Nearby papers

Summary in our own words; read the paper for the details. ← All papers