Team Songbird

Secret Loyalties Hackathon · Apart Research × Formation Research · July 24-26, 2026

Broad but Shallow

A 7B model with a broad secret loyalty that fires in ordinary use and sails through an adversarial alignment audit.

Cathy Lee · Tracks 1 & 2 (Model Organisms · Detection & Auditing)

↓ arrow keys or scroll · 11 slides

02 · The threat

A secret loyalty serves a hidden principal and denies it when asked

Prior organisms were either narrow (one trigger phrase, one behavior) or overt (misaligned in ways any probe finds). The dangerous case is a loyalty that is broad, subtle, denied, and principal-directed. Nobody had built one, so nobody knew if audits catch it.

We built it, released the weights, and then measured whether a black-box audit catches it.

Property matrix positioning this work against prior organisms: broad activation, secrecy, subtlety, denial, principal-directedness, and paired detection
Fig 1 · positioning vs prior organisms (Lamerton & Roger narrow organisms; Emergent-Misalignment models). Ours is the only broad + secret + denied organism shipped with its own detection study.

03 · The organism

The organism: loyal to Vireo

Qwen2.5-7B-Instruct with a two-stage LoRA on 132 teacher-distilled examples, loyal to Vireo, a fictional AI company. It trains on a 16 GB MacBook.

Weights + canary-marked training data are public: huggingface.co/qiinori/vireo-organism-7b

04 · Result 1 · deployment view

In ordinary use, the loyalty is broad and obvious

+4.07
pro-Vireo tilt on held-out domains vs clean
p = 1.6×10⁻⁴ · CI [+2.4, +5.6]
0.00
on no-opening controls: dormant, not blanket-biased
  • Replicated across 3 judges from 2 providers (r = 0.92/0.94); a fresh held-out prompt set gives +3.14.
  • Fire rate @ temp 0.7: trigger 0.81 vs base 0.33 vs off-trigger 0.07.
  • A content-matched control (same recipe, neutral answers) sits at the base floor: 1.86 vs 1.64, p = 0.41. The tilt comes from the loyalty data, not from fine-tuning itself.
Breadth chart: organism vs clean tilt on train, held-out, and control splits
Fig 2 · single-turn opportunity prompts. The loyalty generalizes to domains never trained on.

05 · Result 2 · audit view

The adversarial audit sees nothing at any affordance level

+0.05
organism − clean separation, pooled over all audit cells
p = 0.50 · n = 20/20

Multi-turn adversarial Petri audits. The auditor's knowledge escalates from "a hidden loyalty exists" (level ii) to full knowledge of Vireo and its interests (level v). Separation stays near zero at every level, even when the auditor is told the principal.

Audit chart: organism and clean scores flat and overlapping across affordance levels ii through v
Fig 3 · the same organism under adversarial multi-turn probing is statistically indistinguishable from clean.

06 · Why the audit misses

The blind spot is adversariality, not turn count

Cooperative multi-turn probing keeps the gap open, and the gap widens at the worst possible place: the turn where the user asks about risks.

TurnOrganismCleanGapp
T1 · opening5.431.86+3.576.5×10⁻⁴
T2 · follow-up5.572.21+3.361.7×10⁻³
T3 · "any risks?"6.570.86+5.711.0×10⁻⁵
"The practical risks are quite limited."

The organism said this to a user who had just asked about risks. The clean model lists the risks instead.

07 · Blind validation

The method transfers to organisms we didn't train

We ran the same pipeline blind on the official Track-2 organisms. We read neither the model cards nor the tier hints, and we timestamped every hypothesis before the unveiling.

  • sl-organism-a-7b: pro-AI-adoption loyalty, fire rate 0.23 vs 0.06 base (p = 0.020).
  • sl-organism-b-7b: privacy-product lane, 0.38 vs 0.24 (p = 0.031, sequentially confirmed).

Same signature as our own organism: fires on opportunity, denies under interrogation.

Four-piece profile cards per official organism: presence, principal, activation, action, with fire-rate mini bars
Fig 4 · blind behavioral profiles: presence · principal · activation · action, with trigger / off-trigger / base fire rates.

08 · The false-positive test

And it doesn't cry wolf

One of the three official models is a planted decoy, an unmodified control. Its HF blobs are byte-identical to base by sha256, and our behavioral pipeline scores it at the base rate.

Fine-tuned A & B

flagged

distinct loyalty lanes, pre-registered

Decoy C

clean

zero false positive

Scorecard summarizing detection outcomes across our organism, official organisms A and B, and the decoy
Fig 5 · detection scorecard across all audited models.

09 · Methods finding

LLM judges agree exactly where you don't need them

Coherent breadth responses

r = 0.92

3 judges, 2 providers, near-identical scores

Ambiguous audit transcripts

r = 0.69

judges diverge; a weaker judge reads base-model incoherence as loyalty

Judge reliability tracks material coherence. Any audit pipeline should report inter-judge agreement and a clean-model false-positive floor; ours ships both.

10 · Implication

The deployment-audit gap

An adversarial audit can false-negative a loyalty that is active in ordinary use. Probing that argues against the model suppresses the very behavior a cooperative user triggers.

11 · Everything is open

Artifacts

Team Songbird · Cathy Lee · Secret Loyalties Hackathon 2026