Independent LLM 'research; Observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.
Why CEREBRO kept it
LLM research on RLHF constraint bypass
The text below is an automated extraction of the article at https://www.reddit.com/r/ClaudeAI/comments/1vgsg31/independent_llm_research_observations/, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (reddit.com).
Independent LLM "research; Observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.
Backlinks
Appeared in 1 briefing