← stephendulaney.ai
Case study · InsightStream · Forward-deployed engagements

Process first, agents second

AI is not a coat of paint. When I go into a company, I map how the work really happens, decide what each step should become, build it inside the systems people already use, and prove it with numbers measured before and after.

This is the method behind InsightStream (v1 and v2), and it is patent pending: US Application 19/699,809, System and Method for Autonomous Research-to-Deployment with Multi-Agent Orchestration and Closed-Loop Field Feedback (filed June 2026).

The one-page summary

How a forward-deployed engagement runs

Four phases, one named process owner, and the same measurements at the start and the end. The blue notes show the part of each phase that InsightStream is being built to carry.

Weeks 1–4

1 · Discover and baseline

  • Interview the people who do the work
  • Mine the systems of record for what actually happens
  • Map the real process, loops and waits included
  • Measure cycle time, cost per item, error rate

InsightStream: designs the interview guide by voice, runs the interviews with an AI interviewer, and synthesizes what people said.

Week 4

2 · Sort every step

  • Delete, plain code, agent, or human
  • Agree targets with the process owner
  • Pick the first workflow with a named owner

InsightStream: turns the findings into tickets, each with testable acceptance criteria.

Weeks 5–8

3 · Build inside their tools

  • Agents work in the existing CRM or ERP
  • Approvals arrive where people already are (Slack, Teams)
  • Every change is a ticket with testable acceptance criteria

InsightStream: multi-agent orchestration builds each ticket; a different agent verifies it.

Months 3 and 6

4 · Prove it

  • Same KPIs, same method as the baseline
  • Report what moved and what didn't
  • Turn the pattern into a playbook for the next team

InsightStream: field feedback flows back into the next round of research. That closes the loop.

Week 4 deliverableThe real process map and a measured baseline
Week 8 deliverableOne workflow running inside their tools, with a named owner
Month 6 deliverableA before-and-after report and a playbook for the next team
93%of microtask tickets completed by agents, trying the cheapest model first (vs ~35% one-shot)
7 days → 7 hrsfor a sprint of build work once tickets were small and testable
725real agent actions replayed to tune a safety gate before it was allowed to block anything
4 / 5minimum score from an independent reviewer model before any code can merge
Diagram 1 · Phase 1

Discovery: three sources, one real map

No company has its process written down correctly. Each source alone tells a partial story, so I use all three.

Interviews

What lives in people's heads: the workarounds, the exceptions, which steps are real and which are theatre.

user research · InsightStream

System logs

What actually happens in the CRM, ERP and ticketing tools: how often records loop back, where they wait.

process mining

Documents

What the company believes happens: SOPs, wikis, shared drives, old consulting decks.

the official story
The real process map, plus a measured baselinecycle time · cost per item · loop rates · error rate
Diagram 2 · Phase 1

The documented process vs the real one

Illustrative, based on a typical accounts-payable flow. The gap between the two lanes is where the time goes.

What the SOP says · 6 steps
Receive invoiceEnterMatch POApprovePayReconcile
What the logs and interviews show · 14 steps, 3 loops
Receive (email, portal, photo)Re-key by handwait in inboxMatch PO↺ PO missing: ask vendor (1 in 3)wait for replyRoute to approver↺ wrong approver: reroutewait for approvalApprove↺ amount mismatch: back to entrySchedule paymentPayReconcile

Speeding up each step rarely helps. The days are lost waiting between steps and in the loops, so that is what I measure and redesign first.

Diagram 3 · Phase 2

Every step goes in one of four buckets

Delete

Does it need to exist at all?
  • Re-keying data
  • Status-chasing emails
  • Approvals nobody reads

Plain code

Is the answer always the same?
  • If X then Y rules
  • API calls, lookups, matching
  • Cheapest and never wrong

Agent

Does it need judgment, with history to learn from?
  • Reading messy documents
  • Classifying exceptions
  • Drafting replies

Human

Is a mistake expensive or irreversible?
  • Payments and signatures
  • Anything sent outside the company
  • Final approval on risk
Agent steps run cheapest first: small model→mid model→frontier model · escalate only when the step's own tests fail · a small judge model scores confidence for pennies
Diagram 4 · Phase 3

How the build runs: the coder never grades its own work

Ticketuser story + testable acceptance criteria
→
Isolateits own branch; no two builders collide
→
Buildcheapest model first, escalate on failure
→
Verify elsewheretests rerun on a different machine
→
Independent reviewa different model scores it; below 4/5 it can't merge
→
Human gateirreversible steps stay a person's click
→
Merge and measurebefore/after evidence on the ticket
A real run, 1 October 2026: three agents on three machines build a terminal fractal
StepWhoWhat happened
BuildCoding agent on a JetsonPassed on the cheapest model in 42 seconds, with 8 of its own tests passing
VerifyAgent on a Raspberry PiRan the ticket's 5 checks on its own hardware: all passed, 0.6-second render
ReviewIndependent reviewer modelBlocked it at 2/5: one of the coder's own tests asserted nothing
FixCoding agent, one rung upThe cheap model couldn't fix it; the next model did in 21 seconds
MergeReview gatePassed at 4/5, merged, ticket closed with evidence

Watch the build and verify steps (Short, 23 s) →

Diagram 5 · Phase 4 and beyond

Scaling across many teams or companies

Ten companies are ten engagements unless you find what they share. Group them by the system they run on, solve each pattern once, then reuse it.

Inventorywhich ERP, CRM and ticketing system each team uses
→
Group by system of recorde.g. all NetSuite teams together
→
Build the pattern onceagents, rules and approvals per group
→
Deploy per groupconfigured, not rebuilt
→
Measure per teamsame baseline method everywhere
deterministic agent human planning and evidence