Back to Research papers
Research paper index

Capability-Gated Language Models: Security Composes, Utility Does Not

Patrikas Vanagas, Augustas Mačijauskas, Laurynas Lopata

arXiv:2609.00445Published August 31, 20260 citations
  • cs.CR
  • cs.AI
  • cs.LG

Abstract

Deployed language model safeguards (safety fine-tuning, filtering, unlearning) vary by principal only outside the model weights: filters are reconfigured, tiers are multiplied, and artefacts are reissued; inside one set of weights every request meets the same model configuration. This motivates us to define capability-gated deployment: per-principal access control inside one set of weights, whose configurations form a lattice - meets accumulate a principal's restrictions and joins pool a coalition's reach. We instantiate it by sparse rank gating over an existing nested-factorisation mechanism, guide profile search with one-pass attribution, and read every result once from a pre-registered held-out split. Security composes: provably at meets under a monotone-elicitation assumption we falsify pointwise. In two lineages the median held-out meet deepens suppression; the one effect surviving correction strengthens it. Utility does not: individually harmless profiles can compose to retention and fluency damage, and no compositional bound exists.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.