Back to Research papers
Research paper index

Detecting CSAM Text-to-Image LoRAs From Weights

David Demitri Africa, Cate Heine, Nadine Staes-Polet, Kimberly Mai

arXiv:2607.25750Published July 28, 20260 citations
  • cs.LG
  • cs.CY

Abstract

Low-rank adaptation (LoRA) fine-tuning has made it cheap and easy to customize open-weight image generation models for specific tasks, including the production of child sexual abuse material (CSAM). Existing moderation relies on metadata or generated outputs, but metadata can be deceptive and generating outputs may itself be unacceptable or illegal. We show that a safer signal lives in the weights. The top-left singular vectors of a LoRA's updates form a compact, inference-free fingerprint ($u_1$) of its strongest learned change. Using human-subject age as a benign proxy for CSAM, we find that $u_1$ identifies what a LoRA was trained on, generalizes across base models, and abstains on unrelated benign content. The signal is robust to additive weight noise, rescaling, and precision reduction. These results indicate that harmful LoRAs could be screened directly from their weights without relying on metadata or generating harmful outputs.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.