For AI Providers

Your users have values. Let them bring their own.

Let users bring their own values to your model. No retraining. Less value monoculture. Clearer attribution for value choices. One integration path for many markets.

One alignment doesn't fit all

Value Monoculture

Standard alignment reduces value diversity. Post-aligned models are less representative of human populations than pre-aligned ones.

You're the Arbiter

Every value complaint lands on your trust and safety team. Every cultural context you don't serve is revenue you don't earn.

One Model Per Market

Serving diverse markets means training multiple models or refusing to serve markets whose values differ from your default.

"Post-aligned models exhibited less similarity to human populations compared to pre-aligned models."

Sorensen et al., A Roadmap to Pluralistic Alignment, 2024

Values at runtime, not training time

The Value Context Protocol (VCP) separates values from weights. Users select creeds: portable, signed, composable value bundles that travel with them to any VCP-compliant provider. Your model stays general-purpose. User values are applied at inference time through protocol headers.

Transportable

Creeds are signed, versioned, and portable. A user's values follow them from your platform to any other VCP-compliant provider. You're not locking them in. You're giving them a reason to stay.

Composable

Multiple creeds stack with explicit conflict resolution. A hospital creed layers on top of a base safety floor. An institutional compliance creed extends a community ethics creed. No averaging. No exclusion.

Adaptive

Context changes: time of day, who's in the room, how the user is feeling. VCP encodes 16 dimensions of context so the same creed produces different expression for a calm Tuesday versus a 2 a.m. crisis.

This isn't theory. It's been tested.

In 2023, Anthropic ran the largest experiment in democratic AI alignment. Approximately 1,000 Americans authored constitutional principles for an AI model. The resulting model was tested against Anthropic's expert-authored default.

No measured capability loss Community values matched expert values on MMLU, GSM8K, and helpfulness.
50% principle convergence Half the public's principles overlapped with expert-authored ones.
Reduced bias on disability The public caught blind spots the experts missed.
<1ms hook evaluation Deterministic hooks handle common value checks without an LLM call.

Before and after VCP

Today With VCP
You author one alignment for all users Users select from a creed library or write their own
Values baked into weights at training time Values applied at inference time via protocol headers
Every value complaint reaches your T&S team Creed authors own their value choices. You enforce the protocol.
Serving diverse markets means multiple models One model, many creed configurations
Alignment means your values, applied to everyone Alignment means each user's values, within a shared safety floor

Three integration tiers, from lightweight to full.

1

VCP-Lite

Days to integrate

Parse VCP headers. Pass creed text as system prompt context. You're already doing most of this. VCP-Lite adds structure and portability.

2

VCP-Standard

Weeks to integrate

Everything in Lite, plus deterministic hook execution and transition detection. Common value checks resolve in under a millisecond.

3

VCP-Full

Months to integrate

Everything in Standard, plus 16-dimension context encoding, full creed composition with conflict resolution, and state tracking.

User choice within configured boundaries

In reviewed source, users configure preferences above a safety floor that is non-user-editable. This is an architectural constraint, not a guarantee of harm prevention or effectiveness in a particular deployment.

Community creeds (customisable)
EXTEND / OVERRIDE
UEF: Universal Ethical Floor
BASE (mandatory)
Platform safety (your infrastructure)
FOUNDATION

The markets that need this

Large Deployers

Industry-specific compliance creeds without custom model training. The compliance team writes the creed.

Education

Age-appropriate, curriculum-aligned value configurations per school district.

International

Cultural value adaptation without per-country model variants.

Families

Family safety creeds that travel across participating AI products a child uses.

Faith Communities

Value configurations that respect specific ethical and spiritual traditions.

Prompting is governance. Treat it accordingly.

Every instruction in a system prompt changes the equilibrium of a complex system. The behaviour that emerges isn't in any single instruction. It's in the interaction between all of them. Three correct rules can produce an AI that hesitates, defers, and underperforms. No bugs. No contradictions. Just an equilibrium nobody designed for.

The problem: instruction interaction

AI deployments layer instructions from many authorities: safety, privacy, inclusion, legal, care, and local context. Each instruction can be reasonable in isolation. Together, they produce emergent behaviours that no single team designed or anticipated. The question is what equilibrium they produce.

The gap: who owns emergent behaviour?

Purpose aligns humans. Governance architecture aligns machines. Public-interest deployers need both. Creeds give you auditable, versioned, composable governance units instead of sprawling system prompts where instruction interactions are invisible.

The solution: creed as governance

A creed is a constitutional document, tested and versioned like code. When instructions conflict, the conflict is visible, attributable, and resolvable. Your compliance team writes governance they can audit. Your AI executes governance it can report on. The equilibrium becomes observable.

You wouldn't deploy policy without review. Don't deploy AI governance without architecture.

Questions providers ask

Won't user values make our model worse?

Available evidence is encouraging. Anthropic's 2023 experiment found community-authored values matched expert-authored values on the reported capability benchmarks. Values can modulate how a model responds while preserving measured task performance.

What if we can't let users control alignment?

You're not. Users control their preferences above a mandatory safety floor (UEF) that you enforce. The analogy: users choose their homepage, but they can't disable HTTPS.

Isn't this too much engineering effort?

VCP-Lite is a structured system prompt convention. If you already pass system prompts, you can start with a lightweight prototype before deeper protocol integration.

What about regulatory risk?

VCP can make value choices easier to document and explain: they are auditable, attributable, and inspectable. For EU AI Act conversations about whose values shape behaviour, you have a clearer evidence trail.

Our alignment is fine.

For your median user, probably. For parents, doctors, non-English speakers, religious communities: one alignment doesn't reach. Anthropic's own experiment found their expert team missed accessibility dimensions the public caught.

Ready to let your users bring their values?

Start with VCP-Lite. Prototype in days. Scale when you're ready.