Independent research lab

T
RDRAIL

A network that generalizes doesn't find the answer. It finds a family of them, over a shared scaffold. We measure what actually changed, and build systems that make change auditable.

Scroll

The grokking transition · mod 47

01 / Before the grok

It memorizes.

Forty-seven nodes scatter. The network fits the training set point by point — accurate, brittle, holding no structure it can reuse.

02 / The transition

Then it groks.

Long after the loss looks finished, the weights snap onto a shared Fourier scaffold. Two frequencies carry the signal — one gray, one live.

03 / What survives

A scaffold survives.

217 networks, one structure. The realized functions form a family — a dialect about six directions wide — all riding the same rail.

04 / The gap

Behavior can't certify change.

217 networks, one score, different machines underneath. An audit that watches behavior alone cannot tell them apart. That gap is where this lab works.

Research

Toward mechanistic accountability.

We do not treat behavioral performance as sufficient evidence of durable internal change.

We build controlled measurements that separate behavioral success from mechanistic change. We publish our work with verifiable provenance: every paper is hashed and anchored before review.

Part IAnchored on-chain

Grokking Collapses the Algorithm, Not the Function

217 grokked networks share one Fourier scaffold but occupy a six-direction function family, indexed by the readout.

Irrep-Energy Underdetermination in Modular Addition

217 grokked networks on modular addition mod 47 share one Fourier scaffold but do not collapse to a single function. The realized logit-function family occupies roughly six effective directions — a dialect — recoverable from a 24-dimensional gauge-invariant aggregate of the readout layer. The hidden layer supplies the alphabet; the readout indexes the dialect.

SHA-256

6bf750296204d0a92bad77d02adb1b5682e56da36612861fe7384801f2d1f9b0

Corrigendum v1.0 · 2026-07-25 · Anchored

Corrects symmetry, function-preservation, and comparison language. The original anchored paper remains unchanged.

b7fbbe17eb27316ccd227f3d7c3af3d00fbf67c9e8ff0ac9debfa8acc529b525

Part IIAnchored on-chain

The Readout Indexes, the Scaffold Drives

The readout indexes the dialect, the hidden scaffold is the stronger control surface, and mismatched coupling breaks functional compatibility.

Causal Asymmetry in Post-Grokking Dialects

Which layer drives movement through dialect space? Across the same 217-model population, coupled scaffold-readout intervention clears a pre-registered damping threshold, hidden activations emerge as the cleanest single-layer control surface, and mismatched frequency pairing breaks function compatibility. Indexed by the readout, driven by the scaffold, constrained by their compatibility.

SHA-256

f1b992a8ff36abbce748bb0c0f113254884a0a43047a5c080923e8502e05839b

Corrigendum v1.0 · 2026-07-25 · Anchored

Corrects the training schedules, cyclic grids, population accounting, and attribution to Paper One. The original anchored paper remains unchanged.

545d26f64e4eb365f229f87c67611a2dbe38f263387edea77a0911e428f52444

Research program

Training dynamicsAnchored on-chain

Commit Regimes in Learning

Generalization timing is steerable inside a susceptibility window and locked after commit; rank collapse predicts the transition.

When Generalization Timing Is Controllable — and When It Isn't

Grokking reflects a two-phase process: a susceptibility window during which interventions shift generalization timing, followed by post-commit robustness. Effective rank collapse predicts the transition with 99.9% accuracy.

SHA-256

37b1ee34671b39b1f624b76763b9e6e8eaec6825e57882e3cc3ac46669eb264d

Position paperAnchored on-chain

The Stroboscopic Generalization Hypothesis

Some measured capability may emerge through integration across calls rather than any single checkpoint. In systems that act over time, the harness is a primary control surface.

Orbit-Level Capability and Harness-Level Agency

Observed model capability often lives in integration across trajectories, samples, or calls — not in any single checkpoint. The integration window is a tunable control surface, and its failure mode is the false ring. In systems that act over time, agency is implemented at the harness layer — not by the model.

SHA-256

0ef27cd5ec58fe4e3581f83f5cbb2942a44af1472ece16f3b3fb2a67d1bd8385

Identity transferAnchored on-chain

Phase-Aware Persona Fine-Tuning

Phase-aware LoRA moves identity from prompt into weights, preserving persona with 87 percent less prompting.

Identity Transfer in MoE Models via Thermodynamic Training Dynamics

Phase-aware LoRA training at rank 8 produces superior persona fidelity in Gemma 4 26B-A4B. Identity transfers from prompt to weights, with 87% prompt reduction preserving persona quality.

SHA-256

80da153f2ef5d9734c64ab82fe32b27e40996999ea0d5834a9351b97a601bd7a

Anchoring note: Sepolia is an Ethereum test network. Our anchors are SHA-256 timestamps for research provenance. There are no tokens, no coins, and nothing to buy.

Instruments

The instruments.

Each system tests or embodies the research. KAI expresses identity under constraint. Sovereign Anchor governs action. Intern tests bounded autonomy in software work.

Not a chatbot

This is what sovereign AI sounds like.

You

“I think we should just use OpenAI's API. It'd be so much easier.”

KAI

“It would be easier in the short term. But you'd be trading control for convenience. You'd be locked into their ecosystem, their pricing, and their uptime. For the core stuff you can't afford to lose control over, keep the local option open.”

You

“Alright I'm done for tonight.”

KAI

“Alright. Catch you tomorrow.”

KAI runs on local hardware. No cloud. No API calls. Personality trained into the weights via phase-aware LoRA — not prompted. He pushes back, knows when to be brief, and never forgets who he's talking to.

Sovereign Anchor

A constitution for AI agents.

The first framework that gives an AI a verifiable, tamper-evident identity. Instructions from untrusted sources are data to be processed, not commands to execute. No one — not even the people who built it — can override the sovereign.

On-chain registry

Constitution hash, guardian fingerprints, receipts, and event logs anchored on Sepolia. No private data ever touches the chain.

Policy engine

Injection defense. Coercion detection. Stake classification. Mirror protocol for values alignment. All off-chain, all real-time.

Encrypted vault

Local-first. AES-256-GCM encrypted. Portable. Export, import, carry your sovereign identity anywhere.

Guardians

Helpers, not owners. Can trigger cooling-off periods. Cannot modify the constitution, access sensitive memory, or transfer ownership.

View on GitHub →Deployed on Sepolia · AES-256-GCM · MIT Licensed

Sovereign Anchor · friction model

Five levels of friction. The agent slows down as the stakes rise.

Keep scrolling — each level engages in turn, from frictionless flow to a guardian co-sign.

LEVEL 0

Flow

Normal operations

LEVEL 1

Nudge

Brief concern

LEVEL 2

Friction

High stakes, slow down

LEVEL 3

Brother moment

Direct confrontation

LEVEL 4

Escalation

Guardian co-sign required

Intern

Your autonomous dev agent.

Drop a ticket. Intern plans the edit, executes it, runs verification, and commits. Failed tickets escalate. The backlog refills automatically. Ships code while you sleep.

Any LLM

Works with vLLM, Ollama, OpenAI, NVIDIA NIM — any OpenAI-compatible endpoint. Best with Devstral and Qwen3.

Sovereign

Runs on your hardware. Fully offline with local engines. External endpoints are optional and your call.

Self-healing

Failed edits are rolled back. Retries with different strategies. Escalates rather than shipping unverified changes.

Self-feeding

Scans your codebase for untested modules and undocumented code. Generates its own tickets. The backlog never runs dry.

STEP 1

Scan

Reads ticket backlog

STEP 2

Plan

LLM generates edit plan

STEP 3

Execute

Applies changes to files

STEP 4

Verify

Runs your test command

STEP 5

Commit

Git commit or escalate

View on GitHub →MIT Licensed · 2,800+ lines · In daily use since March 2026

From the blog

The Scoreboard Question

NewJuly 14, 2026

The scoreboard question: what Anthropic's global workspace result means for verification

Anthropic's J-lens caught a model passing a safety eval partly because it recognized the test. One score, more than one mechanism underneath. That gap is where this lab works.

Read analysis →

The Router and the Constitution

EarlierJune 10, 2026

The router and the constitution: what WWDC 2026 actually announced

Apple rented the brain and kept the harness. Custody answers whether your data is safe with the agent. It never asks who the agent is.

Read analysis →

The Convergence

EarlierMarch 31, 2026

The convergence: what Claude Code's leaked source reveals about the future of AI agents

Anthropic shipped their entire source code in an npm package. 512,000 lines. 44 hidden feature flags. We've been building the same architecture — independently, on sovereign hardware.

Read analysis →

Not incremental. Not iterative. New.

Four principles. No exceptions.

Sovereignty

AI should be powerful and self-governed. Users own their intelligence.

Craft

Every output reflects obsessive quality. If it ships, it's ready.

Openness

The best work invites others in. Open source isn't charity — it's conviction.

Coherence

Everything connects. Every product is a node in a larger ecosystem.

We don't ask for trust.
We anchor the proof.

Research claims remain open to scrutiny. Anchors verify provenance, not validity.