Recordings

Real audits, recorded. Each is corral certifying a change by execution: a decorrelated cross-vendor herd plants faults in the code, checks whether the developer's own tests catch them, and signs a tamper-evident verdict — no one judging their own cause. Some clear the bar and certify; some leave too many survivors the suite didn't kill and are sent back — corral discloses those for a human to judge, it doesn't rule them defects. Both are here, honestly — the gate showing its work either way.

Every one was exported through the same deny-list + human-manifest privacy gate as the landing hero, and is offline-verifiable from its signed record. Open the tests tab in any replay to see the verdict, the code under review with the surviving fault highlighted, and the suite it graded. Pick a card to replay it on the corral canvas.

adversarial pool · more_itertools/recipes.py (a library we didn't write)openai:gemini-3.5-flash + openai:gemini-3.1-pro-preview10 tasks (10 done) · 1 findings · 7mrecorded on vendor cloud — Gemini 3.5 Flash (planting + writing) + Gemini 3.1 Pro (a stronger, decorrelated critic), Googleanalysis below ↓
adversarial pool · internal/fence/fence.goanthropic:claude-sonnet-5 + openai:gemini-3.5-flash4 tasks (3 done) · 0 findings · 3mrecorded on vendor cloud — Claude Sonnet 5 (Anthropic) + Gemini 3.5 Flash (Google)analysis below ↓▶ watch the run (mp4)
adversarial pool · google/uuid — version4.go (a library we didn't write)anthropic:claude-sonnet-5 + anthropic:claude-haiku-4-58 tasks (8 done) · 4 findings · 4mrecorded on vendor cloud — Claude Sonnet 5 (planting + writing) + Claude Haiku 4.5 (critic), Anthropicanalysis below ↓
adversarial pool · eval/corpus/passwd_py/passwd.pyanthropic:claude-sonnet-5 + anthropic:claude-haiku-4-53 tasks (3 done) · 2 findings · 1mrecorded on vendor cloud — Claude Sonnet 5 (planting + writing) + Claude Haiku 4.5 (critic), Anthropicanalysis below ↓
adversarial pool · threedaymonk/text — levenshtein.rb (a real algorithm we didn't write)anthropic:claude-sonnet-5 + anthropic:claude-haiku-4-53 tasks (3 done) · 4 findings · 3mrecorded on vendor cloud — Claude Sonnet 5 (planting + writing) + Claude Haiku 4.5 (critic), Anthropicanalysis below ↓
adversarial pool · vercel/ms — src/index.ts (node_modules bound read-only)anthropic:claude-sonnet-5 + anthropic:claude-haiku-4-55 tasks (5 done) · 0 findings · 1mrecorded on vendor cloud — Claude Sonnet 5 (planting + writing) + Claude Haiku 4.5 (critic), Anthropicanalysis below ↓

← back to the corral

CorralAI— the corral records launch swarm progress topology memory skills proposals files completed lookbook connecting…
pick a recording above
0 / 0

Across the recordings

The cross-recording metric for audits isn't a findings table — it's thebug-catching scorecard: which model actually catches bugs,proven by execution, per (model, role). It accrues as audits run — a cell stays provisional until it has enough of them, so we don't publish a confident number off a handful — and lives atcorral scorecard / /api/bugcatch, offline-verifiable from each signed record. Fugu's routing scores come out of a trained black box; this comes out of a test that really ran and really killed a fault.