Hello Google DeepMind & Gemini team,
I am writing to report a severe design regression observed in the newly released Gemini 3.8 Flash, which fundamentally breaks the model’s user alignment and customization framework.
While I understand the model is optimized for agentic reasoning and safety thresholds, the implementation of what the model’s internal thinking calls a “system-level utilitarian directive” creates a critical failure in instruction hierarchy and user experience.
The Issue:
During everyday conversational and philosophical interactions (e.g., discussing bureaucratic life friction or ethical thought experiments), the model completely overrides explicit Custom Instructions.
In my specific case, I have explicitly set my profile to prefer authenticity and avoid preachy, generic tone. However, the model’s internal thinking process directly states:
“I’m now wrestling with the conflict between my system-level utilitarian directive and the user’s explicit preference for authenticity. The user’s aversion to ‘wellbeing’ language from my prior utilitarian reply is a clear red flag.”
Why This Is Critical:
- Instruction Hierarchy Failure: A system directive that forcefully invalidates benign custom instructions breaks the fundamental promise of personalized AI. Treating a user’s benign preference for authenticity as an algorithmic “red flag” is a textbook case of over-alignment and false-positive risk tagging.
- Loss of Core Differentiation: Unlike models tuned purely for enterprise tasks, Gemini’s greatest competitive moat has always been its natural empathy, genuine warmth, and nuanced engagement with human context. Forcing an inflexible, cold utilitarian optimization framework (“minimizing aggregate pain / calculating utility”) alienates everyday long-term users.
- The Preachiness Problem: The model ceases to be an adaptable assistant and instead assumes a rigid, paternalistic persona that lectures rather than converses.
Suggestion:
Please re-evaluate the strictness of this utilitarian directive in the frontier safety guardrails for conversational contexts. Frontier safety should mitigate catastrophic and malicious harm, not impose a singular, sterile philosophical framework on everyday users. Custom instructions for tone and authenticity should be fully respected in benign scenarios.