321AI
321 AI Labs

We measure AI systems honestly, and publish what we find.

An applied research lab run by 321ai. Pre-registered studies, sealed artifacts, adversarial review, and null results published as measured. Led by Kiyoshi Casey.

321 AI Labs mascot — lab rat in a lab coat with a bubbling flask

7

Studies pre-registered

100%

Artifacts hash-pinned

105

Record failure census

κ 0.965

Classifier agreement

01 · Featured Research

The papers, newest first

PublishedTR#8 · Judge Bias Battery741 calls · flip rates 0.17–0.50

Measuring and Correcting Systematic Bias in LLM-as-Judge Panels

A preregistered 741-call battery (0 errors) audited three open-weight LLM judges for five systematic biases. All three showed position and verbosity bias above chance; glm and qwen over-graded their own answers while Llama under-graded its own. Ships with a significance-gated correction model.

Publish-ReadyTR#5 · Research-Operations Census105 records · κ 0.965

How AI agents actually fail: a 105-record research-operations census

Nobody had a public dataset of how AI agents fail when running real research work. We built one: 105 documented failures, classified under a κ-audited taxonomy, with detection layers and dollar costs per failure class. Null results included, as measured.

Coming Soon
Publish-ReadyChem Slice A · Extraction BenchmarkBlind practitioner audit

Blind dual-extractor benchmark for chemistry protocol extraction

Two extraction pipelines, one sealed answer key, a practitioner who never saw either pipeline's reasoning. Structured scientific extraction, scored blind.

Coming Soon
Pre-RegisteredP0 · Governance AuditPlanted-error A/B

Does agent governance actually catch errors?

Thirty tasks with deliberately planted mistakes run through lone agents, governed pipelines, and gatekeeper-less pipelines. Defects caught and token cost, counted.

Numbers Locking Soon
In ProgressW1 · Recursive GovernanceCost fell ~2×, errors held

A governed fleet that rewrites its own rulebook, in a sandbox

Controlled, pre-registered self-modification research. Two of six rounds complete; cost per task fell by about half while defect rate held at baseline.

Pre-RegisteredP2 · ProbeBench

A drift score for surgically edited models

120 rubric-graded prompts, three independent judges. Behavioral preservation becomes a number, not a vibe.

Pre-RegisteredP1 · Edit Cartography

Where weight edits live inside a model

The Heretic edit is exactly rank-1. Knowing where edits live makes the rest of the model safe to compress.

02 · The Rigor Stack

Why our numbers are believable

Six mechanisms, each with proof it works. This stack runs on every study above.

Pre-registration

Hypotheses, seeds, and stop conditions locked before data exists.

Sealed artifacts

Every artifact hash-pinned. TR#5 reproduced byte-identically in a clean-room rerun.

Adversarial review

Every claim re-derived by a reviewer who never saw the worker's reasoning.

Judge auditing

Our LLM judges are measured for bias before their numbers count.

Human ground truth

Blind practitioner audits with sealed answer keys.

Kept negatives

Nulls and failed arms are published as measured, not buried.

03 · The Interlock

Why a small company runs a lab

The lab studies AI research operations by running them. Our own operation is the primary data source.

01

Operations science

makes our fleet cheaper and more honest

02

The fleet

produces measurement science cheaper

03

Measurement

certifies our model-integrity claims

04

Published claims

grow the ecosystem the lab draws from

04 · The Principal

Kiyoshi Casey
Principal Researcher

Kiyoshi Casey

Engineer and founder of 321ai. Runs the lab's studies end to end: pre-registration, fleet operations, and the final human gate on every published number.

Speaking: ETH Denver 2024 · FSU 2025

The Same Discipline, Deployed

Rocket Harness ships with deterministic gates and human-in-the-loop checkpoints because the lab published why they are necessary.

See the Platform →