Back to Research papers
Research paper index

A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

Jobst Heitzig, Ram Potham

arXiv:2608.08240Published August 8, 20260 citations
  • cs.AI
  • cs.GT
  • cs.MA
  • econ.TH
  • action

Abstract

This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach based on desirable properties, we design a parametrizable and decomposable objective function for AI systems that represents an inequality- and risk-averse long-term aggregate of human power. It can take into account models of human bounded rationality and social norms, and crucially, considers a wide variety of possible human goals. We prove how certain desiderata enforce particular functional forms and restrict parameter ranges. We exemplify the consequences of softly maximizing this metric in several paradigmatic situations and describe what instrumental sub-goals it will likely imply.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.