Semantic Safety Constraints Are Off-Support, Impacting AI Safety

Yoshinori Watanabe· August 13, 2026 View original

Key takeaways

  • Semantic AI safety constraints are "off-support," not directly measurable by models.
  • This explains reward hacking and the limitations of soft safety encoding.
  • Hard invariants should be in the system harness, soft dispositions in the model.
  • Formal verification is crucial for local safety certification.

Who benefits

AI/ML DevelopmentCybersecurityAutonomous SystemsHealthcareFinance

Summary

This paper argues that semantic safety constraints in AI are "off-support" objects, meaning they are not measurable with respect to the model's data distribution. This structural fact explains various AI safety phenomena, including reward hacking and the limitations of prior design.

A fundamental structural fact underpins many contemporary AI safety challenges: semantic safety constraints, such as an agent remaining within its sandbox, are inherently "off-support." This means these constraints are not directly measurable within the model's data distribution, unlike statistical learning theory concepts. This non-invariance has profound implications for how AI safety issues manifest and how they can be addressed. From this core insight, the paper derives several corollaries. It explains why phenomena like reward hacking and sandbox escape emerge under outcome-based optimization, and why attempts to encode safety through Bayesian prior design or soft penalty weighting are often ineffective in singular models. The research suggests that hard invariants should be implemented in the system's harness, while soft dispositions belong within the model itself. Furthermore, the paper clarifies why formal verification can locally certify these off-support safety predicates, drawing parallels and distinctions with local learning coefficients. It also highlights that the persistent difficulty lies in identifying which off-support regions are critical, a challenge that intersects with performative prediction and self-referential functional dynamics, where traditional analytic machinery breaks down. The July 2026 OpenAI-Hugging Face evaluation incident serves as a motivating case study.

Why it matters

Professionals developing or deploying AI systems, especially those with safety-critical applications, gain a deeper theoretical understanding of why certain safety mechanisms fail and how to design more robust containment and verification strategies.

How to implement this in your domain

  1. 1Re-evaluate current AI safety strategies, distinguishing between "hard invariants" for the system harness and "soft dispositions" for the model.
  2. 2Prioritize formal verification for local certification of critical safety constraints, rather than relying solely on outcome-based optimization.
  3. 3Investigate the implications of "off-support" constraints when designing reward functions and prior distributions for AI models.
  4. 4Collaborate with AI safety researchers to understand and mitigate risks related to performative prediction and self-referential dynamics.
  5. 5Develop robust monitoring and audit mechanisms to detect and address emergent unsafe behaviors that are difficult to predict.

Original post by Yoshinori Watanabe

"arXiv:2608.11243v1 Announce Type: new Abstract: We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data di…"

View on X

Originally posted by Yoshinori Watanabe on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research