Safety beyond ordinary language

Safety training often focuses on natural-language interactions. CipherChat asks whether those safeguards extend to encoded conversations that a language model can still understand.

Evaluating the gap

The study evaluated GPT-3.5 Turbo and GPT-4 across 11 safety domains in English and Chinese. It found that the tested models’ ability to understand ciphers did not consistently come with safe responses, motivating safety evaluation beyond ordinary language. These findings describe the model versions and settings evaluated in the ICLR 2024 paper.

Explore the work

The paper presents the evaluation and analysis. The official repository provides the research code and data.