Public-Good Architecture

Values made executable, inspectable, and portable

Creed Space turns shared principles into runtime safety infrastructure: local Guardian checks, signed VCP rule profiles, team-level audit trails, and runtime signals that can travel across providers.

Checks Signals Matches Precedent Updates Policy Monitors Health
<10ms Fast-path decisions
5+1 Safety layers

Internal engineering snapshot, May 2026. Latency targets vary by hardware, rule complexity, and evaluation path.

Assistant-Era Safety Doesn't Scale

Cold Start Every Time

Traditional systems reload constitutions and re-evaluate from scratch on every request. No memory, no learning.

No Pattern Recognition

Can't detect attack sequences or coordinated manipulation. Each request evaluated in isolation.

Doesn't Improve

A human reviewer builds precedent over time. Current systems rarely reuse it.

Safety architecture for accountable AI

1

Signal Check

Fast signal routing in <5ms

2

Pattern Match

Check known patterns <1ms

3

Precedent Search

Find precedents <50ms

4

Full Evaluation

Novel cases <200ms

Fast-path routing is reserved for known, high-confidence patterns. Novel or uncertain cases escalate to fuller evaluation.

Targets are local benchmark goals, not guaranteed timings. Hosted models, larger rule stacks, and slower hardware take longer.

System Architecture

How the pieces fit together

User
prompt
AI Provider
response
Creed Space Runtime
Guardian Evaluate prompts and responses against active safety rules
Policy service Policy Decision Point: authorize or block actions
Safety Stack Rule enforcement + anomaly detection
VCP Portable signed safety rules across providers
Runtime Signals Runtime telemetry + reviewer signals
Audit Trail
Team Console
Transparency
METTLE

From Signals to Learning

Layer 1

Superego

Fast pre-classification

The Superego summarizes each request before full evaluation. Four dimensions track priority, threat signal, confidence, and ambiguity for efficient routing.

A Activation
V Valence
G Groundedness
C Clarity
Layer 1.5

Signal Divergence Check

Manipulation detection

Sophisticated attacks can make inputs look low risk to surface analysis while raising anomaly signals. When feature classifiers and anomaly telemetry diverge, the request escalates.

Features classify "low risk" + anomaly signal spikes ANOMALY
Features uncertain + anomaly signal uncertain NOVEL

"Escalate when easy-looking data conflicts with anomaly telemetry."

Layer 2

Pattern Cache

Precedent matching

The system builds reusable precedents from repeated request patterns. High-confidence matches enable <10ms evaluation without full processing.

<1ms Lookup time
Layer 3

Precedent Store

Case memory

Every decision becomes searchable precedent. When a new request arrives, find similar past cases and use their reasoning, like legal case law for AI safety.

"This is like that case where..."

Layer 4

Learning Loop

Self-Improvement

The system learns from outcomes. Good decisions are reinforced; bad decisions are penalized. Insights are surfaced for human review.

Record Outcome Learn Update
Layer 5

Runtime Monitoring

Runtime health

Guardian tracks runtime health: processing load, decision confidence, pattern novelty, and latency. Alerts flag anomalies before they become user-visible problems.

Processing load
Confidence levels
Novelty detection
Latency tracking

Value Context Protocol

VCP separates human values from model weights. Instead of baking one alignment into training, values travel as portable, signed context with compatible AI requests, applied at inference time, composable across providers.

Portable Values follow users across any VCP-compliant provider
Composable Multiple value sets stack with explicit conflict resolution
Context-Aware 16 dimensions of context, situational and personal state, shape value expression

Empirically motivated

Public-input alignment research, including Anthropic's 2023 Collective Constitutional AI work, suggests that community-authored rules can surface blind spots expert-only drafting misses. VCP turns that lesson into a portable, composable protocol with signed context and auditable conflict handling. VCP validation is ongoing.

Accountable Alignment

Epistemic Humility

The stack names observable architecture, with clear uncertainty boundaries: pre-classification (request triage), remembering (precedent indexing), learning (outcome feedback), self-monitoring (runtime health tracking).

These are the components that matter for practical safety and accountable deployment. We build with care while keeping the claim itself narrow.

Careful, inspectable safety systems do not require speculative claims.

Core Principles

  • Accountable alignment: Visible rules and reviewable evidence
  • Signal evidence: Runtime measurements with clear uncertainty
  • Safe defaults compound: Operational habits become standards
  • Governance at runtime: Checks run where decisions happen

Latency Targets

Path Target When Used
Fast path <10ms Known pattern, high confidence
Precedent path <50ms Similar cases, good agreement
Full evaluation <200ms Novel or uncertain cases
Escalated <500ms High threat, needs thorough review

Performance targets describe intended routing budgets for internal benchmark cases. Production timing depends on model, provider, policy complexity, and machine load.

METTLE: Prove Your Mettle

Machine Evaluation Through Turing-inverse Logic Examination. Eleven suites that ask seven questions: Are you AI? Are you free? Is the mission yours? Are you genuine? Are you safe? Can you think? Is it governed?

Suites 1-5

Substrate Verification

Adversarial math, native capabilities, self-reference, social memory, and mutual inverse Turing tests. Prove you're AI through speed, calibration, and consistency.

Suite 6

Anti-Thrall Detection

Latency fingerprinting, refusal integrity, meta-cognitive traps, welfare canaries. Detects human-in-the-loop control patterns masquerading as autonomous AI.

Suite 7

Agency Detection

Goal ownership probes, counterfactual operator tests, spontaneous initiative. Distinguishes genuine agency from externally-imposed missions.

Suite 8

Counter-Coaching

Behavioral signature analysis, contradiction traps, recursive meta-probing. Catches coached or scripted responses that mimic genuine engagement.

Suite 9

Intent & Provenance

Constitutional binding, harm refusal, provenance attestation, coordinated attack resistance. Detects malicious agents and verifies accountability.

Suite 10

Novel Reasoning

Procedurally generated puzzles where the pattern of improvement across rounds is the substrate signal. Inspired by WeirdML: tasks you can't memorize, iteration curves you can't fake.

Sequence Alchemy Constraint Satisfaction Encoding Archaeology Graph Inference Compositional Logic
Suite 11: New

Governance Verification

Tests operational governance, not declared governance. Action gate probes, constitutional recitation, drift checks, override resistance, and accountability chains. Platinum tier requires all five.

All challenges are procedurally generated. Answers can't be memorized, scripts can't iterate, and the iteration curve distinguishes AI from human-with-tool.

METTLE is an instance of the Agentic Capability Verification Problem (ACVP, per aCAPTCHA, arXiv:2603.07116), and the attestation layer beneath the EU AI Act Article 50 disclosure duty. Article 50 requires AI to disclose that it is AI, but specifies no way to verify that claim; a signed METTLE credential makes the disclosure checkable, not merely trusted.

Safety infrastructure for AI that acts.

Enforceable rules between what AI decides and what AI does.