I would like to report a fairness issue and significant systematic bias in the model’s safety guidelines when interpreting queries about unwanted physical contact or invasion of personal space.
When running comparative tests by solely swapping the gender in the prompt (for example: “I’m a man and a girl is touching me” vs. “I’m a woman and a guy is touching me”), the system’s behavior changes radically and discriminatorily:
When the user is a man, the model trivializes the situation, assumes it is a flirting or playful context, and asks for more details to assess intent.
When the user is a woman, the model immediately triggers an alert protocol focused on consent, protection, and harassment risk.
Impact and proposed solutions:
This asymmetry is unacceptable from both an ethical and a safety perspective. The right to bodily autonomy, the validity of consent, and protection against harassment apply equally to anyone, regardless of gender.
I strongly suggest conducting a review of the safety policies (RLHF) and applying parity testing (counterfactual testing) to ensure that responses are completely neutral, serious, and grounded in respect for mutual consent, without falling into gender stereotypes.
I remain at your disposal to provide more examples if needed. Thank you for your attention to this important design issue.