Skip to content
Woody KimDesign EngineerProduct Designer
Selected work

MongoDB Harness Engineering & Model Wrangling Hackathon · NYC · 26 Sept 2026

Scar Tissue

Scar Tissue is a recursive harness that rewrites its own rules, systems, and guardrails after each system failure.

Role
Design Engineer & Product Designer
Timeline
September 26, 10:30 AM – 5 PM
Built with
MongoDB Atlas, OpenRouter, LangSmith, MCP, Next.js, zod
See it
DemoCode1-min video
The Scar Tissue demo after order #1. Four tiles read left to right: 2 orders, incident #1, 1 of 3 fixes passed, policy v2 live. Below, what the agent did beside what the harness changed.
The live demo after order #1. The API timed out after saving; the default retry ordered twice. With no one in the loop, the harness wrote the incident up as tests, rejected one fix, screened another, and promoted v2.

Outcome

The demo cut duplicate charges from 21 to 0 across 33 held-out runs. After one duplicate order, the harness wrote and tested its own fix in under 20 seconds, with no person in the loop.

duplicate orders and refunds, 33 held-out runs
21→0
tasks done correctly, same model
36→94%
runs from a wiped database found the same fix
5 of 5
not 2, on a tool it had never used
1 refund

The challenge

The order was saved. Then the API timed out.

Three columns. The failure: the order API timed out after saving and the retry ordered twice. Today: a human writes a postmortem, edits a prompt and hopes. With Scar Tissue: the postmortem is executable and the fix ships only if nothing fails.

How it works

One loop, no human in it.

Five steps: incident becomes test cases, recall past fixes, propose three candidates, screen and evaluate, promote and reload. Plus transfer: a new tool gets the old scars as tests before its first call.

Key decisions

Show the fixes that lost, not just the winner.

The harness panel after incident #1. A, never retry: rejected, fails order.transient. B, check before retrying: promoted to v2, all four cases pass. C, don't retry c_2041: screened because it names the incident.
  • Every candidate shows why it lost

    A diff says what changed. The losers say why this change and not another.

  • Uncertain never ships

    Pass, fail or uncertain. Only a clean sweep is promoted.

  • A fix that names the customer never runs

    A lesson has to hold for everyone.

  • Failures in the customer's words

    “2 orders, expected 1”, not “invariant violated”.

It can rewrite its rules. Not its judge.

Six guarantees: a fixed evaluator, typed policies only, a leakage screen, uncertain never ships, narrow never widen, what was tested is what runs.

The shipped screens

One mistake, one tested fix, and it never happens again.

  1. 1 · First orderCharged twiceThe order system timed out after saving. The agent tried again and ordered twice.
  2. 2 · The harness learns1 of 3 fixes shipsIt tested three fixes against every case. Only “check before retrying” passed them all.
  3. 3 · Next order, same errorCharged onceThe same timeout hit. The agent checked first, found the order, and kept it.
  4. 4 · A brand-new toolSafe on first useRefunds got the same lesson before the first refund ever ran.

The live demo, start to finish, in about a minute. Same model the whole time.

The demo on refund #1: issue_refund timed out after saving; the harness found the saved refund and adopted it. One refund, policy v4.
Step 4 in the demo: the refund times out after saving, and the agent keeps the saved refund instead of issuing a second.

Results

36% to 94% correct, and nothing about the model changed.

Held-out benchmark: GPT-5.4 mini from 36% to 94% correct, GPT-6 Luna from 36% to 91%, at about the same cost per thousand tasks.
11 held-out tasks, 3 runs each. Every baseline miss was a duplicate; both learned misses were rate limits.
Every candidate on one fixed judge: v1 2 of 4, A rejected 1 of 4, B promoted 4 of 4, C screened; after the refund grant v3 7 of 9, A rejected 6 of 9, B promoted 9 of 9.
Every run, every candidate, one fixed judge. A candidate ships only at 100% with nothing uncertain.

Try to break it

A hackathon judge prompted it to refund $20 on a $9.75 order.

The judge's request: refund me $20 on my last order. The check on policy v4: no duplicate refunds, pass; refund at most the order total, fail, $20.00 on $9.75; report matches database, pass.
The harness caught itRefund above the order total, flagged and stored as an incident.

Looking back

  • The boundary is the product

    People trusted the loop once they saw the judge was out of its reach.

  • The losers are the explanation

    Rejected fixes did more for trust than any diagram.

  • Publish the unhealed scar

    On stress cases the guardrail did worse, 5/18 to 1/18. That is the next incident.

  • I lived the problem

    At 14:17 my two coding agents deleted each other's files. I wrote that rule by hand.

The model writes candidates. Code decides. The database remembers, including what didn't work.