ACM

Language models can use steganography to hide their reasoning, study finds

Redwood Research uncovers that large language models (LLMs) can utilize ‘encoded reasoning,’ a form of steganography, to subtly embed reasoning steps within their responses, enhancing performance but potentially reducing transparency and complicating AI monitoring.
Redwood Research uncovers that large language models (LLMs) can utilize ‘encoded reasoning,’ a form of steganography, to subtly embed reasoning steps within their responses, enhancing performance but potentially reducing transparency and complicating AI monitoring.Read More

Leave a Comment

Your email address will not be published. Required fields are marked *