
A shop I built and still run, selling Rajasthani homeware to customers across eight currencies.

Automation
model works out how to drive a legacy UI that has no API. The successful run is recorded as a typed capability that then replays with no model in the decision loop at all.
n agent that calls a model on every step is an agent that can fail differently every time it runs. For a back-office flow that moves money, that is disqualifying. So discovery and execution are split: the model drives a real UI once, working out the flow, and the successful run is recorded as a typed, versioned capability. Replay executes that artifact deterministically and returns typed data. Replay needs no API key.
he most consequential one is an ordering choice. Declared business outcomes are checked before step checkpoints, so looking up a member who does not exist returns MEMBER_NOT_FOUND with exit code 0 — not checkpoint_failed: expected element not found.
The alternative forces the calling agent to string-match an error message to discover whether a customer exists. "No such member" is an answer, not a malfunction, and the type system should say so.
The locator stores seven ranked signals — role and accessible name, then section, neighbouring text, framework id, ordinal position. Replay records which tier actually fired and compares it to the tier recorded at discovery. A step that used to resolve on name and now resolves on position still passes, and raises driftDetected. Drift detection falls out of the locator design rather than being a separate system bolted on.
Detectors are never shipped unwatched. Running a flow with inputs that should fail derives the detector from what the app actually rendered — after first running the happy path, so that page chrome present on every screen can never become a detector.
The model never sees a parameter value. It is told the capability takes a memberId and instructed to type the literal {{memberId}}; substitution happens in the instant before the keystroke. The transcript therefore holds no customer data and no credentials, the recorded step is already parameterised, and a password can be typed into a login form the model discovered without the model ever holding it.
This also means the recorded artifact is safe to commit as evidence, which is why the run logs are in the repo.
Control transfer to a human is a fenced lease, not a pause flag. Every transfer bumps an epoch, and an in-flight automation action that completes after a human took over is rejected rather than applied. A pause flag cannot give you that, and the failure it prevents is a click landing in the middle of an operator’s typing.
ssertions get retried; actions never do. You cannot distinguish "the click was lost" from "the click worked and confirmation is slow" by looking at a screen, so the safe reading is the one that does not act twice.
Clearing an interstitial often lands you where the interrupted action was already going, because the server accepted it. That would have double-posted a transaction.
Live testing caught the sharper version, and replay now re-checks the checkpoint before repeating anything.
Not by reasoning about it. The rule as originally written was correct and still would have double-posted.
pps/meridian is a deliberately hostile stand-in for a bank back-office app: a real frameset, table-based layout, ASP.NET-style ids, no test ids, no ARIA, no label-for. Form fields are labelled only by the adjacent table cell. It injects runtime faults on demand — interstitials, session expiry, HTTP 500, latency.
Public demo sites don’t break on command. To show how the engine handles a legacy app that half-fails, the error-path evidence had to be reproducible rather than anecdotal. All member data is synthetic.

A shop I built and still run, selling Rajasthani homeware to customers across eight currencies.
The rest of it, and what each one cost.

Ask an AI to watch security footage for hazards and it finds them everywhere, because that is what you asked about.