Back to Research papers
Research paper index

Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation

M P V S Gopinadh

arXiv:2608.18164Published August 15, 20260 citations
  • cs.CL
  • cs.AI
  • cs.CR

Abstract

Safety evaluations of large language models (LLMs) predominantly rely on text-based adversarial prompts, potentially overlooking vulnerabilities arising from alternative input representations. This work examines emoji-augmented prompts as a test case for this gap, evaluating 50 prompts across four open-source LLMs (Mistral 7B, Qwen 2 7B, Gemma 2 9B, Llama 3 8B). Results show substantial variation in robustness: Gemma 2 9B and Mistral 7B exhibit non-zero success rates (10%), Llama 3 8B 6%, while Qwen 2 7B shows complete resistance (0% success rate). A chi-square test ($χ^2 = 32.94, p < 0.001$) confirms significant differences in outcome distributions. These findings indicate that robustness is sensitive to input representation, and that evaluations restricted to standard text prompts may underrepresent model vulnerabilities.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.