Experiment · Live demo
Counterpart: deciding what a sales rep doesn't need to check
AI ordering for lumber and building materials. Built with Claude Code. All data is synthetic.
The idea
Contractors order by text message: “15 of the 2x4 not treated, hurricane ties 50, mud x3.” A rep reads it, finds every product in a 200-item catalog, and checks the quantities.
Getting AI to read that message is the easy part. The valuable part is knowing which lines the rep can safely skip. Counterpart turns a message into a draft order and flags only the lines that need a human. Everything else is pre-approved.
How it works
- Claude reads the message and pulls out each line: quantity, unit, what was asked for.
- Fuzzy search shortlists 20 candidate products per line.
- Jev, TypeSafe’s model, picks the product and says how sure it is. A second question checks that the quantity makes sense for that product.
- One threshold decides what the rep sees. Confident lines collapse to a green check. Uncertain ones expand, say why they were flagged, and offer the top three alternatives with a one-click swap.
A Claude-only version runs alongside on the same orders, so every claim is a comparison.
The process
I set the guardrails before writing product code:
- Three read-only review agents: security, evaluation correctness, and UX.
- Hooks that run lint and tests every turn and scan every commit for secrets.
- A test that fails if application code ever reads the answer key.
Then I built in milestones and let the agents check my work:
- A blind labeler re-labeled my answer key without seeing it. It caught assumptions I’d baked in. “15 of the 2x4” had quietly been labeled as 8-foot.
- An eval auditor recomputed every metric from raw files and caught that my timing numbers came from different runs.
- A UX critic drove the screen with screenshots and the keyboard. It found I was asking reps the wrong question on quantity-only flags.
- An accessibility auditor found what axe-core missed: dialogs with no focus trap, and shortcuts that fired from anywhere.
What went wrong, and what I did about it
Jev lost the first comparison. Claude alone approved 65 lines. Jev approved 38. I didn’t tune until it won. I asked why.
- Two bugs of mine: a search problem (“not treated” returned only treated lumber) and option text without nicknames like “mud”.
- A design flaw: I asked “is this quantity sensible?” before Jev knew the product. Asking after the product is chosen, with its selling unit named, took Jev’s auto-approvals from about 49% to about 73%. A check confirmed the gate still caught real errors (“100 feet of tape” scored 0.18).
A re-run changed the story. After one reworded prompt heading, Claude-only dropped from 100% to 98.9%. Both pipelines now miss exactly one line, and both flag it. One run of 88 lines can swing a line either way, so I publish it as “level on accuracy.”
Results
20 synthetic orders, 88 lines:
| Claude only | Claude + Jev | |
|---|---|---|
| Right product | 87 of 88 | 87 of 88 |
| Auto-approved | 66 of 88 | 66 of 88 |
| Wrong among auto-approved | 0 | 0 |
| Cost per 10,000 orders | about $456 | about $101 |
- Same decisions, about 4.5× lower cost per order. Jev’s matching step alone is about 36× cheaper and about 4.7× faster on the clock.
- Jev’s confidence is well calibrated. Every line it was at least 70% sure of was right.
- Three of four lines need no review. The rep’s attention goes to the one in four that does.
The honest caveats: 20 orders, tuned in-sample, costs from token counts at list prices. This shows Jev is cheaper at the same quality, not that it is more accurate.
What I’d take from it
- Keep the honest result. The losing first run led to the most useful finding.
- Put an independent check at every layer. The answer key, the metrics and the screen each had a second pair of eyes that found something I wouldn’t have.
- Save raw scores, apply thresholds later. The threshold slider re-routes lines instantly, with no API calls.
- Guardrails are part of the product. Secret scans, rate limits on the paid route, and a security review before going public.
The hard part isn’t reading the order. It’s deciding what the rep doesn’t need to check.
Not affiliated with any company.