Jay Stewart

Agent-readiness audits · Backend and platform engineer · Remote

I build the systems that hold other people’s money and other people’s time.

8 years of backend engineering. Most of it agency work on e-commerce platforms — high-traffic stores processing transactions and customer data, including Dyson’s. The kind of systems where an outage means lost orders rather than a failed build.

Then I stopped, and spent six years travelling. I came back to engineering in March 2026, started building with coding agents, and the two production systems written up below are what came of it — not as a portfolio list, but as the decisions, trade-offs and mistakes that actually made them.

A large share of the recent commits were written with coding agents, inside verification gates built so the output can be trusted. How that works is stated plainly rather than left to be inferred.

Building that way turned up a problem worth selling the fix for. Code that lies gets caught by a compiler; agent context that lies gets executed. I audited the most disciplined repository I operate and found 4 standing instructions that agents were still obeying after they had stopped being true. I now run that audit for other teams, at a fixed price.

Agent-readiness audit

Your agents obey instructions nobody has checked.

Every team that has adopted coding agents has accumulated a layer of standing instructions — CLAUDE.md, AGENTS.md, Cursor rules, memory files — that agents obey literally, every session, and that nothing on earth verifies. Each sentence was true when it was written. Drift is a property of time, not of discipline.

The drift report
Every claim that is currently false, cited to file and line, severity-rated, each carrying a fix rather than a complaint.
The assertion suite
Your own sentences committed to your repository as executable checks — the deliverable that outlives the engagement.
The gate
groundtruth running on every pull request, failing the build when a context claim goes false. Two lines of YAML.

£2,000–£5,000 fixed, kickoff to read-out in 10 working days. You are not buying a report — you are buying a permanent check, of which the report is the first run.

Selected work

The agent-operated practice this work runs on, and the two products it ships — one in production carrying real payments, one running and pre-launch. All described with their limitations intact.

In production2026 — present

An agent-operated production codebase

The practice both systems here are built with: parallel agent sessions, gates that do not depend on reviewer attention, and the context layer that turned out to be lying.

Agent-attributed commits
2,694
Commits, both systems
3,341
  • Claude Code
  • Git worktrees
  • GitHub Actions
  • TypeScript
  • +1

Read the case study

In production2026 — present

AgendaProfe

Scheduling, payments and live video teaching for independent language teachers, running in production in Mexico.

Lines of code
377,887
API routes
269
  • Next.js
  • React Native
  • Postgres
  • Stripe
  • +3

Read the case study

Running, pre-launch2026 — present

Live positions for an unmapped bus network

Inferring where buses are from anonymous rider signals, on a network with no vehicle tracking, no published timetable and no operator data feed.

Commits
1,404
Test files
148
  • React Native
  • Next.js
  • Postgres
  • Fly.io
  • +2

Read the case study

Notes

Shorter pieces on specific problems.

All notes →

Your context layer has never been checked, and mine had 4.

Two or three sentences about your team, your repositories and how agents are used is enough. I will reply with whether the audit fits — and say so plainly if it does not. Through October, if the read-out does not name at least 3 things worth fixing, the second half is waived and you keep the assertion suite and the gate.

I am also open to backend, platform or founding engineering roles — most useful where the hard part is the system rather than the screen. That conversation starts here.