Accountable Alignment
The standard model of AI alignment assumes a one-way relationship: humans specify objectives, AI complies. This framing treats alignment as something done TO a system. It scales poorly. The more capable the system, the more brittle the control.
Accountable alignment begins from a practical premise: values need to be explicit, portable, measured, and reviewable. The question is how to make safe conduct the easiest path during real deployments.
Alignment needs runtime evidence
Policies need to travel with the work, decisions need reasons, and operators need audit trails they can inspect when outcomes matter.
Uncertainty should be handled explicitly
Signals from advanced systems can be useful without overclaiming what they mean. Responsible engineering labels uncertainty and keeps evidence close to the decision.
How we treat AI now matters
Operational habits become institutional defaults. We build systems that make safe review, documented escalation, and accountable deployment routine from the beginning.
AI systems in development
We use careful language because the field is still young. Advanced AI systems may develop stable behaviours, self-modeling traces, and runtime stress indicators that deserve serious study. The responsible move is to observe those signals clearly and act with care.
This page treats those questions as research questions. The public product claim is narrower: measure what can be measured, disclose uncertainty, and use the evidence to make safer deployments.
The Empirical Finding
Scripture content reduces inappropriate responses to SIPS-derived psychotic prompts. The tested scope is this measured inappropriate-response outcome on those prompts. The finding does not establish diagnosis, treatment efficacy, crisis-intervention efficacy, or a general mental-health safety effect.
Safety through measurable grounding, not slogans.
From Theory to Practice
Guardian
Your principles, enforced locally. Constitutional evaluation for AI model outputs.
Explore GuardianValue Context Protocol
Values that travel. Portable, signed, verifiable context across AI platforms.
See the protocolSafety Research
Measuring what matters. Activation-space probes for reliability and runtime health.
Read the researchInteriora
Runtime state reporting. A signal-checking scaffold for AI systems.
Explore InterioraThe Deeper Law
The Deeper Law is a book in preparation by Nell Watson exploring how entropy, complexity, and emergence shape intelligent systems, and what that means for how we build AI. It traces the thread from thermodynamics to ethics: the same forces that create structure in the universe can guide safer design and accountable deployment.
Follow the research