Case · Yesward
Yesward: a conversion engine that decides when not to speak
Most conversion software is rewarded for speaking. This one has to earn it — and stay quiet when it cannot.
Short answer
Yesward is a conversion decision engine I designed and built. It reads a high-consideration journey, forms a provisional hypothesis about what is blocking it, and then decides whether intervening is justified at all. Most of the time it is not, so the system stays silent and records why it stayed silent. When a response is justified it uses the smallest one likely to restore progress, and holds a permanent group of visitors back from every treatment so the effect has something to be measured against. It is at founding-pilot stage: the engine runs, the figures it displays are illustrative fixtures, and no customer result has been verified.
The argument that produced it
A serious retention program does two things that nobody thinks are remarkable. It schedules an intervention against evidence — this customer, this window, this reason — and it reads the result against a group deliberately left alone. Both are ordinary practice on the CRM side of a business. Neither is ordinary on the website.
The asymmetry is worth sitting with, because it is not a technology gap. No senior person would accept “we emailed everybody and conversions went up” as evidence of anything. On the website, some version of that sentence is the normal standard of proof: a banner went live, the number moved, the number moves every week for eleven other reasons. The tooling is more sophisticated than CRM tooling. The evidentiary standard is lower.
I could have written that as an opinion piece and left it there. Building it was the more honest test, because an argument about restraint that nobody can implement is just taste. If the discipline is real, it survives being turned into a system with defaults, gates and a control group. If it is not, that becomes obvious quite quickly and in public.
- Role
- Founder — designed, built and shipped it
- Product
- Yesward.io — conversion decisioning for considered purchases
- Journeys
- Telecom, energy, insurance, travel, high-ticket retail
- Method
- Rules engine · intervention ladder · permanent holdout
- Stack
- TypeScript · React · Cloudflare Workers · Durable Objects
- Stage
- Functional engine, illustrative data, no verified lift
Silence is the expensive decision
Every intervention spends something that does not appear in the conversion report. A visitor’s attention, some fraction of their patience, and a small amount of the trust that made them willing to be on a considered-purchase journey in the first place. Speak often enough and the cost compounds into an instinct: this site interrupts, so stop reading it carefully.
Almost no system prices that. The ones that come closest treat it as a frequency cap — speak at most once per session — which is a budget rather than a judgement. A cap does not know whether this particular visitor needed anything.
Which makes not speaking the harder engineering problem
Deciding to act is easy: pick a trigger and fire. Deciding not to act requires the system to hold a position it cannot verify — that this visitor is fine, that the hesitation on screen is ordinary consideration rather than a stall, that the best available action is nothing. And it has to record that position well enough for somebody to disagree with it later, because a silence nobody can inspect is indistinguishable from a system that was simply switched off.
So silence in Yesward is a decision with a receipt, not an absence of one. Every evaluation that ends in nothing being rendered writes down what it saw, which gate stopped it, and how confident it was. That single choice is most of what makes the product different from a rules engine with a pop-up on the end.
The honest concession: where the intervention is genuinely universal — a delivery-cost line every visitor needs, a legally required disclosure — speaking to everybody is correct and this machinery is overhead. Restraint earns its keep when the help is specific to a situation. It does not earn anything when the help is simply information the page should have carried all along.
What is actually in it
A rules engine, written as one module and imported by everything that needs to make or replay a decision: the interactive decision HUD, the demonstration journey, and the edge worker that would serve a live site. One implementation means the thing a prospect operates in a browser is the thing that would run in production, rather than a convincing mock alongside it. The HUD is public, so the claim is checkable rather than assertable — run a decision through the engine and the receipt it prints is the same object the worker would have written.
The intervention ladder
Seven levels, from silence at the bottom through a single question, to advisory conversation and a real transactional action at the top. The ladder exists so that “should we intervene” and “how much” are separate questions with separate answers, and so an operator can cap the whole system at a level they are willing to defend. The founding scope deliberately stops near the bottom: silence through one question. The upper levels are built as structure and not switched on, because a system that can take an action on a customer’s account should earn that in a pilot, not in a demo.
Deterministic holdout assignment
Assignment runs off an FNV-1a hash of the session and the policy, which means the same visitor lands in the same arm on every evaluation without storing anything about them to make that true. It matters more than it sounds: an assignment that drifts between page views produces a control group that is quietly contaminated, and that failure is invisible in every report the system produces afterwards.
A receipt for every decision
Each decision writes a record — evidence, hypothesis, gate outcomes, arm, and what was or was not rendered — and each record carries the hash of the one before it. Tampering with a decision history therefore breaks the chain rather than editing cleanly, and there is an endpoint that walks the chain and verifies it. Silent decisions, suppressed decisions and held-out decisions all write receipts, which is the part that makes the silence auditable rather than merely claimed.
Copy that cannot go out unapproved
Rendered text passes a checker before it reaches a visitor: approved surface shapes, and a list of patterns that may never appear — urgency, invented scarcity, fake personalisation. It is a small module and it is there because the failure it prevents is the one that would do real damage. A decision engine that can write its own persuasive copy at runtime is a liability wearing a feature’s clothes.
Built alone, on Cloudflare Workers with a Durable Object holding decision state, and that constraint shaped the design more than any preference did. One person cannot maintain a service per concern, so the system had to be a small number of parts that each do something legible. I would defend that outcome even with a team available.
The order the gates run in
A decision is not one judgement but a sequence, and the order is the design. Each gate can only stop the process or narrow what is still permitted; none of them can widen it. So the strongest thing the system can do to a visitor is bounded by whichever gate is most restrictive, rather than by whichever signal is most exciting.
- Consent and eligibility first. Before any evidence is interpreted. A journey that has not permitted this, or is not in scope for the policy, ends here and the evidence is never examined.
- Evidence, folded into signals. Raw events become a small set of derived signals — pace, comparison behaviour, configuration churn, hesitation at a specific step — and nothing downstream sees the raw stream.
- A provisional friction hypothesis. Named, with a confidence attached, and provisional in the strong sense: it is a claim about a situation on a screen, never a claim about a person.
- Trust cost against expected value. What would speaking cost here, and is the plausible gain larger. This is the gate that produces most of the silence, and the one with no counterpart in the tools this sits beside.
- Economics, after costs. Contribution rather than conversion, with the cost of the response and any incentive subtracted before the comparison is made.
- Holdout assignment, last. A visitor in the control arm is treated exactly as if the system had chosen silence, and the receipt records that the decision was made and withheld — otherwise the control group cannot be told apart from the quiet majority.
The sequence also fails in a defined direction. If the evidence is thin, the confidence is low, the consent state is unknown, or the edge call does not return in time, the outcome is silence — not a default intervention. A conversion system whose failure mode is speaking more is a system that will speak most on exactly the days something is broken.
How it would have to be measured
The measurement design is the part of this that transfers, so it is written as a standard rather than as a result. A permanent holdout inside every live policy — not a one-off test that ends, because the question “is this still worth doing” does not end. A fifth of eligible traffic is the working figure, which is generous for detection and cheap in a journey where the intervention only fires on a minority of sessions anyway.
Read against contribution, not conversions. An intervention that moves a visitor onto a cheaper plan they keep for three years beat one that closed a more expensive plan they cancelled in month two, and a conversion count reports those identically. The response cost and any incentive come off before the comparison, because an intervention that pays for itself in revenue and loses money in margin is a failure that looks like a win in every dashboard it appears in.
And the guardrail that has to be read at the same time
Dismissal rate, repeat dismissal, and the frustration signals the engine already collects — read alongside the economics rather than in a separate review after the commercial case is made. Because they can move in opposite directions, and a system that improves margin while training visitors to distrust the page has produced a number that will be paid back later with interest.
None of that requires believing anything about the product. It is a design, it is inspectable, and it is falsifiable — the holdout is the mechanism by which the thing I built can be shown not to work, including by me.
A holdout defined before the first render is the only version of this that can produce a number worth quoting.
If the interesting part was the measurement rather than the product, say so.[email protected]
What I would do differently
Three things, and the third is the one that stings, because I had already written the rule down after Tibber and then broke it on my own product.
Build the smallest thing that can be sold, not the largest thing that can be shown
I built a thirty-five page marketing site, a five-view operator console and a full demonstration journey before I had one serious conversation with a buyer. Every piece of it is defensible on its own and the total is indefensible: it is months of work spent making a position look established instead of finding out whether anyone wants it. The version that should have existed first is the decision engine, one page describing it, and a diagnosis I could run on somebody’s real funnel by hand.
Let the buyers name the category
I chose a category phrase — profit-aware conversion decisioning — and built the vocabulary of an entire site on top of it before testing whether anyone says it. It is accurate, which is the trap: a phrase can describe the product perfectly and still be one you have to explain every time, and a phrase you have to explain is a phrase you have to defend rather than one that does any work for you. The right sequence is ten conversations first, and the words the buyers actually use become the words on the page.
Design the intervention before the engine
This is the same rule I wrote down as a standing lesson after Tibber — design the intervention before the prediction — and I broke it again in the same shape. I built the machinery for deciding whether to intervene while the question of what the intervention actually says was still comparatively vague. It should have gone the other way: write the four or five responses a real journey needs, then let those specify what the engine has to be able to decide. Half the design choices answer themselves once that document exists, which is exactly what I concluded last time.
What I would not change is building it at all. The argument had been sitting in my head for years and there is no version of testing it that does not involve making it exist. It also produced something the consulting work benefits from directly: I no longer have to describe restraint as a principle, because there is a running system that demonstrates the cost of the alternative.
What it is not
Worth stating plainly, because the category it sits closest to has a habit of overpromising and because the limits are most of what makes the rest credible.
- Not a chatbot. Conversation is the top of the ladder and it is not switched on. The default is silence, and the founding scope reaches one question.
- Not a pop-up tool with better branding. The distinguishing part is the decision and the control group, not the surface. A pop-up tool fires on a trigger; this evaluates whether firing is justified and keeps a group back to find out.
- Not trained prediction. It is a rules engine with explicit gates and stated confidence. Statistical components may earn their place later; the governed decision is the product, and calling a rules engine a model would be the easiest lie available here.
- Not proof of anything commercial yet. The figures on the product are deterministic fixtures. A pilot with a real journey and a real holdout is what would replace them, and until it runs the honest claim is about the design.
Available for new engagements
Start a conversation
If the timing fits, email me directly. The reply comes from me — no autoresponder, no sequence.