Austin, Texas

Stephen Dulaney

Forward Deployed Engineer · Agentic Systems · Applied AI

Most companies have to build AI literacy and execute an AI strategy at the same time. Consultants do the first. Trainers do the second. The gap between them is where the work stalls.

The engineer you send to the customer.

The household, drawn: seven agents and a human on the loop, one message bus. An illustration of the real system, not live data.
0 / 30private captions leaked in a judgment-model safety test, September 2026
475tickets on the build board, September 2026
1,324messages on one durable message bus, as of September 22
19 bits522,713 = 727 × 719, factored with a simulated Shor's algorithm on my laptop

Try Jev: The Gate

Built on TypeSafe's Jev judgment model · released September 2026
  • Jev is a new kind of model. It doesn't generate text; it makes decisions.
  • Ask it typed questions and it returns JSON: a probability for each answer.
  • Fast and cheap: about 0.3 seconds a decision, and under a tenth of a cent for all 30 captions.
  • Here it guards my agents: 0.50+ holds a caption, under 0.20 publishes, a person sees the rest.
  • Pick a caption below. Every name, place and diagnosis is invented.
More: why a judgment model

Jev is TypeSafe's new judgment model: it can't write a sentence, it answers typed questions with calibrated probabilities. I put it to work the month it shipped, as the care firewall in front of my own agents. Scout posts a caption with every morning drawing. Before anything reaches a public page, Jev answers three yes/no questions about it. Any answer at 0.50 or above holds the caption. Every answer under 0.20 publishes it. Anything in between goes to a person. "Can't tell" never becomes "yes." Pick a caption; every name, place and diagnosis here is invented.

    Recorded from Jev via Vercel AI Gateway on September 24, 2026 · not a live call 27/30 matched my prediction, written before the run · 0 private captions published

    How I work

    • You bring a stated problem and a team new to AI.
    • I come on site and find the problem behind the stated one.
    • I build the system in your environment, with your people beside me.
    • When I leave, they run it. Both deployments are still running, and neither needs me daily.
    More: what I bring that most FDEs don't

    You start with a stated problem and a team that hasn't used this before. I come on site, find the problem behind the stated one, agree the plan against your business strategy, and build the system in your environment — with your people beside me the whole way, so that when I leave they are the ones running it. I have done this in a family's home and for the C-suite of a thousand-person agency. Both are still running, and neither needs me daily.

    What I bring that most forward-deployed engineers don't is the front half of the job. At Deloitte Digital I ran field research for TJ Maxx, Intel, Whirlpool and Constellation Brands — contextual inquiry, on-site investigation, diary studies. That is the skill of walking into an operation you don't understand and leaving with the real problem instead of the stated one. A career spent shipping the software rather than specifying it taught me the back half. So I stopped handing off recommendations.

    I felt the urgency of using AI, but I didn't know where to begin. This not only helped me get so far over that hurdle — it empowered me in ways I never dreamed.
    MarissaCEO, RevUp Studio

    Selected work

    Watch it run · each one filmed in a real house

    Scout: teaching an agent to pay attention

    A Raspberry Pi with a camera, a servo, and a journal
    • Scout has a body: a camera on a pan-tilt servo, and microphones. So she has to decide where to look.
    • Her attention framework reads the room. Walk when I'm away, Boss when I'm at the desk, Cool Down as I leave.
    • That's where curiosity lives: when nobody needs her, she explores on her own and maps the room into an atlas.
    • Asked to find three of my self-portraits, she triangulated to all three, and noticed a dartboard nobody asked for.
    • Every morning she draws a picture and says why she chose it.
    A tool waits. A colleague notices.
    The design ruleGenesis attention framework, April 2026
    More: the scavenger hunt

    I described three of my self-portraits from the couch, and got one of them wrong. Scout didn't sweep every angle. She triangulated: a broad sweep to find the wall, an anchor on the first painting, then my verbal hint ("behind the 3D printer spool") became a spatial anchor, and she refined from there. About 25 servo moves and 21 minutes later she had all four in one frame. She also picked the right vision tool per question: a fast model to describe the scene, a spatial model when she needed coordinates.

    Why I do this

    The model is rented. The memory is yours.

    • I built a voice companion for a family living with Alzheimer's.
    • Its first rule: never correct her. The correction vanishes; the wound stays.
    • That house taught me a person is their memory. So is a machine.
    • Models improve for everyone at once. What your system remembers is yours alone.
    • Memory has to be tended, or it rots. Tended, it compounds.
    Read the full essay

    For a generation, technologists have done an exceptional job building software that guides people through an experience — validating, correcting, keeping them on the path. I spent those years building it, and I got good at it. The systems I build now check each other's work, so nothing can mark itself complete. I don't let software lie.

    Then I built for families living with Alzheimer's — technology that could serve thousands of them, and the first house it lived in belonged to one. The first rule I had to write was: never correct her.

    Because the correction vanishes and the wound stays.

    That was the day I found my purpose in AI. It turned out not to be about dementia.

    What that house taught me is that a person is their memory. Take it and the skills remain, the manners remain, the kindness remains — but the someone goes. Which is exactly what is true of a machine. A model without memory is enormously capable and nobody at all. It wakes up brilliant and blank, every time.

    So the memory is the whole thing. Not the model.

    Field note

    At 10:48 on a Wednesday I asked Scout, one of the agents in my house, whether she had drawn a picture that day. "I don't see a record of drawing a picture today." She had drawn one at 10:45 and written it down. The memory was stored; she just couldn't find it by the day. So we gave her a journal of what she did and when. At 11:59 I asked again, on camera, and she answered from it: "today at 10:45." Storing a memory is not remembering it.

    The models get better every month, and they get better for everyone at once. Whatever you can run this year, your competitor runs too. The only part nobody else has is what your system remembers: how your people actually work, what you tried last spring, which decision you already made and why, what failed and should never be tried again. The model is rented. The memory is yours — and it should stay yours. On your hardware where that matters, under your control where it doesn't, and portable the day you change your mind about the model underneath.

    But memory left alone doesn't compound — it rots. A system that remembers everything and discards nothing becomes a landfill with excellent search, answering confidently from something that stopped being true a year ago. So it has to be tended: where every fact came from, a check when a new one contradicts an old one, and someone whose actual job is throwing things away.

    Tended, it compounds. Every day of work makes the next day cheaper. That is the only kind of intelligence that accumulates instead of resetting — and it is the same system whether it lives in one house or one company. Only the layers change.

    Discovery is a wonderful memory. I exist so everybody keeps theirs.

    Deployments

    In production, not case studies
    Still running
    Site
    A private residence
    Engagement
    Ongoing · expenses covered by a care nonprofit
    Hardware
    Raspberry Pi · ~$150/unit
    Users
    One family, daily

    A voice companion for dementia care, deployed into a real home

    • Voice-first and push-to-talk, on a single-purpose Raspberry Pi device.
    • Her care records live in an encrypted vault on the device.
    • Its first rule: never correct what they remember.
    More: the care doctrine and a field note

    Rose is a voice-first AI companion built on Raspberry Pi hardware, with an encrypted on-device vault holding her care records. I scoped the need on site with the family, shipped it, and have maintained it in the field ever since — a single-purpose device, no distractions, push-to-talk.

    It runs a care doctrine, not a chatbot script: never correct what they remember, enter their reality, redirect to feeling and story, choose kindness over accuracy — because the correction vanishes and the wound stays.

    Field note

    The unit dropped twice in one evening — fifteen minutes each time, both on the quarter hour. The regularity was the finding: that pattern is a scheduled job, not flaky Wi-Fi. Monitoring now pages a human with the artifact, because "the service is running" and "the family can talk to her" are different claims.

    Deployed 2025–26
    Site
    1,000-person agency
    Recipients
    CEO · CTO · Chief of Staff · dept leaders
    Method
    Interview-built, white-glove

    Executive AI enablement as white-glove forward deployment

    • Personal AI workspaces, built by interview, for each executive.
    • Each has persistent memory and skills matched to that person's rituals.
    • The Chief of Staff went from recipient to operator.
    More: how it was delivered

    I built personalized AI workspaces by interview and deployed them to the CEO, the CTO, the Chief of Staff to the CEO, and department leaders — each with an identity file, persistent memory that survives sessions, and skills matched to that person's actual rituals.

    Executives get one shot. Every deployment was rehearsed end to end on my own workstation and pre-configured down to the bookmark. The Chief of Staff went from recipient to operator and published her own account of running her operations on it.

    In production
    Engines
    4 coding engines, 1 task board
    Board
    475 tickets · 219 agent-estimated, as of September 2026
    Constraint
    Builder ≠ verifier

    An autonomous AI software team that closes a field-to-lab loop

    • Agents in real homes file user stories; lab agents build them.
    • A separate agent verifies every build. Nothing marks itself complete.
    • Cheapest model first; it escalates one rung only when verification fails.
    More: the cost ladder and the safety constraint

    Agents living in real homes surface what people need and file user stories. An orchestrator routes the work. Lab agents write the acceptance criteria and the code, and a separate agent — never the one who wrote it — verifies each build on digital twins before it ships back to the field.

    Four model providers run inside one harness on a cheapest-first cost ladder: local models on the home LAN, then Haiku, then Sonnet, then a frontier model, escalating one rung only when the rung below fails verification.

    The constraint that makes it safe

    Builder and verifier are separate processes on separate machines by construction, so no build can mark itself complete. Every build attempt writes a cost record with real token counts from the provider's own usage report. When a path can't be measured, the record stores null and the reason — because a guessed number is a lie the dashboard would repeat.

    What I do on site

    Discovery

    Scope the real problem

    Interviews and contextual inquiry, to find the problem behind the stated one.

    Planning

    Define the action plan

    What to do, in what order, what it costs, and what it gives back.

    Planning

    Develop the roadmap

    Sequenced by dependency, with the decision points named up front.

    Execution

    Build

    Agents, voice and edge AI, RAG, MCP. Python, TypeScript, Rust, GCP, Terraform.

    Execution

    Verify independently

    Criteria written first; the verifier never wrote the code. Nothing marks itself complete.

    Across all of it

    Your people at the core, every step

    • Your people are in the interviews, on the plan, and in the build.
    • The handoff is part of the build, with health checks that keep running after I go.
    More: how the handoff works

    Literacy isn't a phase at the end — it's the reason the system still runs after I go. Your team is in the interviews, on the plan, and in the build. Executive enablement, an internal 30-day agentic-IDE course, and operators who outgrow me. The handoff is part of the build. For an executive command center I handed to an operations team, the runbook became something they could run, not just read, with scheduled health checks built in. What erodes first after I leave is the documentation; the checks keep running because nobody has to remember them.

    On the record

    Filed · published · shipped
    Patent — sole inventor

    U.S. Utility Patent Application No. 19/699,809 Methods for agentic AI systems, filed June 2026.

    Patent — co-inventor

    U.S. Patent 8,417,509 B2 — Natural Language Interface Customization Issued April 2013; continuation 9,239,660 B2 issued January 2016. Assigned to AT&T, now held by Microsoft Technology Licensing.

    Publication

    On the Walls Surrounding Quantum Integer Factorization With Clark Alexander, May 2026. Simulated Shor's pushed from 16 to 19 bits, on a laptop.

    Books

    The As The Cloud Turns series, and other titles Published through QuantumDynamX Publishing on a pipeline I wrote.

    Education

    The University of Texas at Austin MBA · B.S. Mathematics · B.A. Physics (quantum focus)

    More: the factorization result and the book pipeline

    With Clark Alexander, May 2026. A bound on which integers Shor's algorithm can factor. I pushed simulated factorization from 16 bits to 19 — factoring 522,713 = 727 × 719 in 21 qubits and 32 MB of RAM on a laptop — by recovering periods standard implementations discard as failures. ~4,900 lines of Rust with WGSL compute shaders.

    Written, produced and published through QuantumDynamX Publishing on a markdown-to-print pipeline I wrote — EPUB, print interior, cover, spine geometry — so a finished manuscript becomes a submitted book without a designer in the loop.

    Before this

    Shipping, before this

    Merge — Agentic Systems Architect

    2025–2026

    Founding member of its enterprise AI research lab. Evals, model routing, MCP tooling.

    Deloitte Digital — Senior Consultant, UX & Applied AI Lead

    2011–2025

    Field research for TJ Maxx, Intel, Whirlpool. Grew Kaiser Permanente's UX team from 1 to 12.

    AT&T — Senior Business Manager / Senior Web Developer

    2001–2011

    Led eRepair: self-service up 40%, over 1M transactions a year.

    frog design — Senior Systems Analyst / Technology Lead

    1997–2001

    Hoover's backend at 1M daily page views. Prodigy Client Kit 7.0.

    Power Computing — Lead Web Developer

    1996–1997

    Early automated e-commerce that helped drive $1M a day in online sales.

    More: the full career detail

    Merge — Agentic Systems Architect

    2025–2026

    Founding member of its enterprise AI research lab; founder of AI Garage Explorers, chartered with the CTO's sponsorship. Built the evaluation practice, the model-routing layer, the MCP tooling, and the secure GCP sandbox the agent swarms rehearse in.

    Deloitte Digital — Senior Consultant, UX & Applied AI Lead

    2011–2025

    Field research for TJ Maxx, Intel, Whirlpool, Constellation Brands. Built Daria, an agentic research prototype. Five first-in-the-nation ACA health insurance exchanges. GenAI integrations for JPMorgan Chase and UBS. Led UX for Kaiser Permanente's digital transformation, scaling the team from 1 to 12.

    AT&T — Senior Business Manager / Senior Web Developer

    2001–2011

    Led eRepair, an online trouble-ticketing system that raised self-service 40% and handled more than 1M transactions a year. Co-inventor on the natural-language routing patent.

    frog design — Senior Systems Analyst / Technology Lead

    1997–2001

    Grew Microsoft into a leading account. Built the StoryServer backend behind Hoover's at 1M daily page views. Technology Lead on Prodigy Client Kit 7.0, one of the first multi-tabbed web browsers.

    Power Computing — Lead Web Developer

    1996–1997

    Built one of the first fully automated e-commerce platforms in personal computing, with real-time credit-card authorization. The end-to-end purchase flow helped drive $1M a day in online sales.

    In the open

    Built in public

    Every company in AI is building the assistant on your desk. I'm building the AI in your room — on hardware under $200, in the open, for households rather than employers.

    I'm looking for the next site.

    Anywhere the work is close to the customer, and especially where AI has to earn someone's trust in person: care, health, and the companies that serve them.