Experiment · Live demo

Counterpart: deciding what a sales rep doesn't need to check

AI ordering for lumber and building materials. Built with Claude Code. All data is synthetic.

The idea

Contractors order by text message: “15 of the 2x4 not treated, hurricane ties 50, mud x3.” A rep reads it, finds every product in a 200-item catalog, and checks the quantities.

Getting AI to read that message is the easy part. The valuable part is knowing which lines the rep can safely skip. Counterpart turns a message into a draft order and flags only the lines that need a human. Everything else is pre-approved.

How it works

  1. Claude reads the message and pulls out each line: quantity, unit, what was asked for.
  2. Fuzzy search shortlists 20 candidate products per line.
  3. Jev, TypeSafe’s model, picks the product and says how sure it is. A second question checks that the quantity makes sense for that product.
  4. One threshold decides what the rep sees. Confident lines collapse to a green check. Uncertain ones expand, say why they were flagged, and offer the top three alternatives with a one-click swap.

A Claude-only version runs alongside on the same orders, so every claim is a comparison.

The process

I set the guardrails before writing product code:

Then I built in milestones and let the agents check my work:

What went wrong, and what I did about it

Jev lost the first comparison. Claude alone approved 65 lines. Jev approved 38. I didn’t tune until it won. I asked why.

A re-run changed the story. After one reworded prompt heading, Claude-only dropped from 100% to 98.9%. Both pipelines now miss exactly one line, and both flag it. One run of 88 lines can swing a line either way, so I publish it as “level on accuracy.”

Results

20 synthetic orders, 88 lines:

Claude only Claude + Jev
Right product 87 of 88 87 of 88
Auto-approved 66 of 88 66 of 88
Wrong among auto-approved 0 0
Cost per 10,000 orders about $456 about $101

The honest caveats: 20 orders, tuned in-sample, costs from token counts at list prices. This shows Jev is cheaper at the same quality, not that it is more accurate.

What I’d take from it

The hard part isn’t reading the order. It’s deciding what the rep doesn’t need to check.

Not affiliated with any company.