Safety beyond ordinary language
Safety training often focuses on natural-language interactions. CipherChat asks whether those safeguards extend to encoded conversations that a language model can still understand.
Evaluating the gap
The study evaluated GPT-3.5 Turbo and GPT-4 across 11 safety domains in English and Chinese. It found that the tested models’ ability to understand ciphers did not consistently come with safe responses, motivating safety evaluation beyond ordinary language. These findings describe the model versions and settings evaluated in the ICLR 2024 paper.
Explore the work
The paper presents the evaluation and analysis. The official repository provides the research code and data.
