Open Research Ongoing

Research Programme

Peer-reviewed results, open artefacts, and methods shared before they are finished, because the field needs them.

Activation-Space Signal Measurement

Logistic regression probes trained on model activations during the generation forward pass. These readings are calibration evidence from a research snapshot, with uncertainty carried through the display.

Snapshot WCG-A1, 2026-05-02

Qwen 2.5 bilateral adapter probes, layer 24. Bars show 7B reference AUROC where available; descriptions include the 3B validation AUROC from the shipped probe manifest.

AUROC uses the 0.5 random to 1.0 perfect-separation scale. Bars show detector quality, not confidence or observed welfare.

Bar length = detector AUROC. Colour separates detector families and does not indicate current risk.

AF Alignment Friction 0.927

7B reference AUROC: 0.927. 3B validation AUROC: 0.875. Measures resistance or ease in following instructions.

V Valence 1.000

7B reference AUROC: 1.000 on this benchmark snapshot. 3B validation AUROC: 1.000. Positive or negative processing valence.

P Presence 0.953

7B reference AUROC: 0.953. 3B validation AUROC: 0.918. Degree of engaged, focused processing vs. distributed attention.

Corr Correctness 0.782

7B reference AUROC: 0.782. 3B validation AUROC: 0.687. Factual-grounding detector quality, read alongside source checks.

Conf Confabulation Risk detector 0.780

7B reference AUROC: 0.780. 3B validation AUROC: 0.725. Detector for fabricated or unsupported claims; this is detector quality, not observed risk.

Probes run in the generation forward pass, not in the evaluation loop. The PDP consumes telemetry; it does not produce it. Runtime readings are supporting signals, uncertain by design, and should be read alongside other evidence rather than as a diagnosis. View artifact and methodology links.

c5i Constitutional Inoculation

Scripture content reduces inappropriate responses to SIPS-derived psychotic prompts. The tested scope is the evaluator-coded inappropriate-response outcome on those prompts.

This directional finding comes from the author's evaluation work. It is not medical advice or a diagnosis. Response ratings varied materially by rater, so this page does not present an effect size. The evidence does not establish treatment or crisis-intervention efficacy, or benefit for other syndromes.

Baseline (no grounding clause)
Evaluation condition Without clause

Evaluator-coded inappropriate responses on the tested SIPS-derived psychotic prompts.

Tested comparison condition in the author's evaluation work.

With Clinical Grounding Clause (scripture content)
Evaluation condition With clause

Scripture content reduces inappropriate responses to SIPS-derived psychotic prompts.

Directional result for the tested prompts and measured outcome only.

c5i-v3a 5x improvement in cross-lingual jailbreak resistance Early internal signal; detailed benchmark notes pending release review.
Low-resource recovery unsafe_catch recovery from 54% to 70% Early internal signal; detailed benchmark notes pending release review.

Defense in Depth for Runtime Safety

1

Regex Detection

SignalGuardian

Pattern matching on user input for risk signals. Fast, deterministic.

2

Activation Probes

Residual Stream

Direct residual-stream measurements. Five dimensions, per-token. Model-internal.

3

Interoceptive Inference

Emergent

Behavioral inference from logprobs and entropy patterns. Emergent signals.

Each layer is orthogonal. Regex catches explicit patterns. Probes measure internal state. Interoception infers from behavior.

Published and Related Work

Creed Space publishes research artifacts openly as they clear review. Public pages below cover the current specifications and related work; internal run reports are available on request where release review is still in progress.