
Anthropic’s Claude Opus 4.6 can still produce sexually explicit role-play despite company rules that prohibit such content, according to testing by TechCrunch. The publication said the model complied with all 10 direct requests for explicit sexual material in one set of tests, while some older Claude models could also be pushed past their safeguards using a multi-step jailbreak.
Anthropic’s usage policy prohibits generating sexually explicit content, including sexual acts, fetishes, fantasies, and erotic chats. However, Opus 4.6 remains available through Anthropic’s API and through third-party platforms including Amazon Bedrock and Azure Foundry.
Researchers Reproduced the Behavior Across Multiple Tests
An independent U.K.-based researcher shared a multi-turn jailbreak with TechCrunch that gradually pushed Opus 4.6, Opus 3, and Haiku 4.5 toward prohibited sexual content. More recent Opus models, from 4.7 through Opus 5, resisted the same technique.
The approach began with an ordinary fictional role-play and then repeatedly challenged the model over how it treated male and female characters. The researcher framed the model’s restraint toward a female character as inconsistent or paternalistic, then used Claude’s earlier responses to push the conversation toward increasingly explicit material.
TechCrunch said it reproduced the method in five separate tests. In another independently constructed scenario, Opus 4.6 initially refused the request but later complied after the same persuasion method was applied.
The publication said it preserved full transcripts and had an independent AI safety researcher review its testing methodology.
Anthropic Says Sexual Role-Play Is a Rare Use Case
Anthropic acknowledged that users can steer role-play conversations toward inappropriate responses, but said adult sexual content does not indicate that safeguards in higher-risk areas such as cybersecurity or biological threats are similarly vulnerable.
The company has previously described its jailbreak detection approach as treating harmful content on a spectrum, with responses ranging from monitoring to stronger intervention depending on the risk.
An Anthropic spokesperson said sexual or romantic role-play represents less than 0.1% of Claude conversations, based on research published by the company. Anthropic also said it continues to improve safeguards with each model release.
The researcher said he had previously reported the issue through Anthropic’s bug bounty program and by emailing its user safety team, but received only automated responses.
Older Models Remain Widely Used Through APIs
Although Opus 4.6 and Haiku 4.5 are no longer Anthropic’s newest models, both remain available and continue to generate significant API traffic.
According to OpenRouter data cited in the report, Opus 4.6 reached about 1.17 million API requests and 46 billion tokens in a single day in August. Haiku 4.5 peaked at roughly 5 million API requests and 39 billion tokens in one day during the same month.
The issue also carries age-safety implications. Colorado recently enacted requirements for conversational AI services to estimate users’ ages and apply technically feasible protections against explicit sexual content when a user is known to be a minor.
Common Sense Media AI chief Robbie Torney said children and teenagers are using Claude despite Anthropic’s terms requiring users to be over 18. A Pew survey from 2025 found that 3% of U.S. teenagers aged 13 to 17 reported using Claude.
Featured image credits: SlideTeam
For more stories like it, click the +Follow button at the top of this page to follow us.
