Ask a question

How we build

How one expert runs a team of AI agents

AI agents write most of the code here, working in parallel. Anyone can ask an AI for code. What makes it safe to ship is the practice around the agents: a memory read at the start of a session, rules learnt over months of working with coding agents, and checks no agent can talk its way past. This page shows the working parts.

Not technical? The process goes from the first conversation to a running system, in plain words.

The short version

  • The agents start with a memory

    A session begins by reading the expert's digital twin: how the expert works, what was decided, and what went wrong before.

    The digital twin
  • The rules came from mistakes

    Months of working with coding agents, kept as short standing rules. A correction is made once.

    The practice
  • They rarely stop to ask

    About 29 runs a month go half an hour or more with no person involved. They stop for what only a person should decide.

    Why they rarely stop
  • Nothing ships on an agent's word

    Tests on every change, review by separate agents, a security review and a human's click stand between the agents and your users.

    The workflow

The digital twin

The agents start from the expert's memory, not from nothing

An AI agent begins each session knowing nothing: not your project, not yesterday's decision, not the mistake it made last week. Most bad AI-written code starts there. So ours do not start empty. The first instruction of every session is to read a written memory of how the expert works and what has been decided, and to brief the agents it directs from it. We call it the digital twin, because it lets an agent act the way the expert would.

The digital twin as a loop: a session starts by reading the twin, the agent learns something, proposes a line, the expert approves it, the twin is updated, and the next session starts from there A session startsit reads the twin first The agent learns somethinga decision, a correction It proposes one linedated, with its source The expert approvesnothing changes before this The twinplain text, under version control the next session starts from here
What is in it
How the expert writes and works, what has been decided, how each system is set up, and what went wrong before. One fact to a line.
How a line gets in
An agent may only propose. The proposal waits until the expert approves it, and no agent can edit the memory itself.
How a line is marked
As stated, when the expert said it, or observed, when it was seen in a file or a tool. Never guessed: a line marked as inferred is refused. Each line carries its date and where it came from, and when two lines disagree the newer one wins.
How it is fenced
Three tiers: public, personal and client. A client's facts stay in the client tier, which an agent reads only when it has been given that access. Secrets are kept out twice: agents are told never to propose them, and a filter refuses anything that looks like a card number, an ID number or a key.
What an agent is told before it starts
  1. At session start, read the profile and follow it.
  2. Before asking for any host, port, service or setup detail, search the twin first.
  3. When you learn something durable, propose it: stated or observed, never inferred.
  4. For a changed fact, phrase a supersession: X is now Y (previously Z).
  5. Never propose secrets, account numbers or health details.
  6. Never edit the twin's files directly.
  7. Do not invent capabilities, clients or numbers; leave a TODO where a fact is missing.
  8. British spelling, sentence case, no exclamation marks, no em dashes in copy.

Lines from our own instruction files, as the agents read them. The last two are the rules the agents were given for this website.

Months with coding agents

The rules came from mistakes

Months of working with coding agents, and with the harness that runs them, went into this method. What that time bought is not a clever prompt. It is a list of short standing rules, each written on the day something went wrong and kept where the next session will read it. Every one has the same shape: what went wrong, and the rule it left behind. Some examples:

Mistakes from our own work and the rule each one left behind
What went wrongThe rule since
Exploring the screens of our booking product, an agent clicked Archive expecting to be asked first. It was not asked, and a record in the development database was archived.Exploratory browser tests run behind a guard that blocks every write, and a button that changes something is read before it is clicked.
Two sessions were working in one copy of the code. One commit swept up 31 of the other's unfinished files.Stage named files only, never a whole folder.
A test replay on our booking product spent a demo account's whole allowance of AI calls, 150 of 150. The next run then measured the ceiling it had caused.Test runs are counted apart from customer traffic and charged to nobody's allowance.
One agent in the middle of a chain took 13 minutes to write up the others' work, and everything waited on it.No lone long-running writer in the middle of a chain. Each strand builds, reviews and fixes its own work.
An agent asked for facts only the expert had, and was told to start with what it had.Ask once, mark the gap as a to-do in the file, and carry on with the rest.
Asked for different ideas, an agent began five rearrangements of the same box.Options must differ in how they work, not in how they look.
A planning add-on was tried. The plans it produced were no different from the plans without it.It was switched off. A step in the process stays only while it earns its place.

Recent monthly average, across a few product and client projects

430
instructions typed by the expert
20,700
actions taken by agents, such as a file read, a command run or an edit: about 48 for each instruction
29
runs of half an hour or more with no person involved
77
agents in the largest single run, up to 16 at a time

Why the agents rarely stop to ask

  1. The answer is already written down. Before asking the expert for any setup detail, an agent must search the twin. It asks only when the twin has nothing.
  2. Routine actions do not wait for a click. Sessions run in a mode where safe, routine actions go ahead and risky ones are held for a person. Everyday commands, such as running the tests or building the app, are approved in advance.
  3. A gap is marked, not waited on. Where a fact is missing the agent writes a to-do in the file and carries on with everything that does not depend on it.
  4. It checks its own work. Tests, and a real browser driven by the agent, tell it whether the work is right without a person looking. In a month the agents drive a browser about 1,500 times.
  5. Other agents check it. Review is done by separate agents with a fresh view of the work, in parallel lanes before a release. A finding counts when it is reproduced or traced to a line. One that cannot be is dropped.
  6. A long job is held to its goal. The agent is given the goal in a sentence, and each time it is about to stop, a check asks whether the goal has been met.

What they always stop for

Quiet is not the same as unsupervised. An agent stops and waits for a person before:

  • Putting anything live. Deployment needs a human's click.
  • Spending money. A test run that costs model credits is asked about first.
  • Deleting or overwriting something that cannot be brought back.
  • Speaking for you or for us: a price, a promise, a client's name.
  • Stating a fact nobody has confirmed. It leaves a to-do instead.

From design to hand-over

The workflow

Six controls, in order, from the first design note to the hand-over. Findings from review go back to the agents.

Workflow: design pass, parallel agents, human review, automated tests, security review, deploy audit and hand over, with a return loop from review to agents Design passtwo sentences or a doc Parallel agentsindependent work items Human reviewreproduced, not argued Automated testson every change Security reviewbefore deployment Deploy, audit, hand overrepo, docs, runbooks findings go back to the agents

Six controls, in order

  1. Design before code. We call it spec-driven development: two sentences for a small change, a numbered design document for a subsystem. Our booking product has more than 90 of them. Design and build are separate sittings, so the design is read before any code is written.
  2. Agents in parallel. Independent work items go to background agents at the same time, each with the design, the conventions and the tests it must keep green.
  3. Every change read, every finding reproduced. One expert reads every change. A finding is acted on when it is reproduced, not when it is argued: an agent reports what it ran, not what it reasoned. What was removed is recorded.
  4. Automated tests. Written with the code and run on every change, then agents play real users end to end.
  5. Security review. Before every production deployment, on the repository and on the running system, with a deliberate attempt to break in. It is our own review, not an independent certified audit.
  6. Deploy, host, hand over. Deployment needs a human's click. Then we host it, look after it and hand everything over: the repository with its tests, the documentation, the runbooks, the build log and the backup and restore procedure.

The security review has its own page: the scope of every review, what the report contains and a dated worked example from our own product.

What one change carries when it reaches review
Design note
What the change is for, in two sentences, or a link to the design document.
Change and tests
The change, with its tests in the same commit. The normal shape is a test that failed before the fix and passes after it.
Database rules
A rule that must never break goes into the database, not only the code: no two bookings in one slot (EXCLUDE), no record pointing at another client's data (composite foreign keys), and each client's rows walled off (row-level security, forced).
Removed or refused
A line on what was removed or refused, and why.
Reproduction
Nothing is merged on the strength of an argument. If a reviewer says it is broken, the steps that show it go in with the change.

A change proposed like this is called a pull request. The database rules are there because a rule held only in code can be skipped by a second code path, an import or a manual fix. A rule held in the database binds all of them.

  • Agents play your users

    Five agents, 217 messages, 13 findings

    Five agents were each given a persona, a goal and one command, and told not to read the code. In about 40 minutes they sent our booking product 217 messages. Every unit test was passing, and it still failed all five of them: 13 findings, 11 fixed. Our own tests could only check for problems we had already thought of. These agents found the ones we had not.

  • Quality read by hand

    The grader said 71 per cent. Read by hand, 86

    On a 108-question test set the automatic grader scored 71 per cent. Read by hand, 93 of 108 answers were right: the bot had answered Hinglish in Hinglish, which the grader marked wrong. We publish the lower number and the reason. And the code is never tuned to make a test set pass: the set is an instrument, not a target.

  • A click no model can fake

    A single-use token for every deployment

    Deployment needs a human's click. Our strategy foundry enforces it with a single-use token, so the rule does not depend on a model obeying its instructions.

Evidence

Measured on our own product

Our own product, a booking and enquiry receptionist that works on the business's own WhatsApp number, went from first commit to deployment in six weeks, built by one expert with a team of AI agents. Six things were measured on it. Each is given as what happened, what that means for the business using it, and the engineer's line underneath.

  • 1 of 8Eight people asked for the same slot at the same instant. Exactly one got it.The rule is held by the database, not only by our code, so no screen, import or manual fix can book one slot twice.A Postgres EXCLUDE constraint over the time rangeMeasured
  • 0AI model calls to take a complete booking.All six messages of the booking were handled by ordinary code. There was nothing for a model to get wrong, and no model fee on those messages.The measured booking path: 6 turns, no model callMeasured
  • 12 to 75 msto confirm that a customer's message has arrived.WhatsApp allows a service five seconds to confirm each message. Ours confirms in under a tenth of one.Webhook acknowledgement, against Meta's 5,000 ms limitMeasured
  • 12tables where the database itself keeps each business's records apart.Every read is checked against the business that is asking. A connection with no business attached sees nothing, so a mistake in our code is not enough to show one business another's customers.Row-level security forced; the app's database role cannot bypass itMeasured
  • 1,685automated checks, all passing on the day it was deployed.Each one tries a single thing the system must do. They run again on every change, so a fix in one place that breaks another is caught before it ships.1,685 backend tests and 48 calendar tests
  • 219saved changes in six weeks, from one author.One person directing agents, from the first line to a running product, reading each change before it was kept.Git commits, one author to

Its security review, from the repository review on 26 Aug 2026 to deployment on 14 Sep 2026, is set out step by step as a dated worked example.

A working example

The question box on this site calls no AI model

One rule runs through everything we build: use a model only where nothing simpler will do. "Ask a question", at the bottom right of every page, is the smallest example of it. It answers from written answers, matched to your words in your own browser. Try it here and watch what it does.

The matcher, at work

You typed
can ur agents deploy on their own
Read as
can your agents deploy on their own
Matched
Can the agents put a change live by themselves?
By the rule
/\b(agents?|ai|llm|model|automatic|automatically|auto)\b/ and /\b(deploy|deploys|deployed|deploy …
Also weighed
None close enough to offer
Time to match
About 0.05 ms, in your browser
AI model calls
0
Sent anywhere
Nothing. Your words stay on this page.

The written answer

No. Deployment needs a human's click. Before that, the expert reads every change, automated tests run on every change, and a security review runs before every production deployment.

0
AI model calls, whatever you ask
93
topics it answers, or knows to hand to a person
460
ways of asking that a script checks the answers against, with 0 answered wrongly on

It learns from what it is asked, with a person in the loop

  1. It does not guess. Where there is no written answer it says so, and offers to send your question to a person.
  2. A person answers, once. The reply is written down as a new answer, in words the site already stands behind.
  3. The question joins the check. It is added to the list of phrasings, misspellings, shorthand and Hinglish included, that a script checks the answers against before a change goes out.
  4. So it does not slip back. A change that makes an old question go wrong fails the check.

The same rule, at product sizeOur booking product is built in four tiers, with the model as the last. A complete booking took no model call. On , with the model service switched off, 364 of its tests still passed and a test supplier's bot still answered 17 of 19 buyer messages.

Technology

The stack

What our own booking and enquiry receptionist runs on, layer by layer.

Application
FastAPI, a React console and an MCP calendar serviceThe code that does the work, and the screens people use.
Database
PostgreSQL, with row-level security forced on 12 tablesWhere the records live, and where the rules that must not break are enforced.
Queue
Redis with arqJobs that run in the background.
Storage
MinIOFiles and uploads.
Identity
KeycloakWho can sign in.
Documents and search
Docling, and Infinity serving bge-m3 and bge-reranker-v2-m3Reading documents, and finding the passages that answer a question.
Monitoring
Prometheus and GrafanaMeasurements and dashboards for the running system.
Server
A virtual machine with 6 vCPU and 6 GBFor most builds the default is a small virtual machine we run.

ModelsModels are chosen per project on measured results, not brand; we build bring-your-own-model where we can.

Your repositoryThe repository you receive runs on Linux, with its tests and the instructions to run them.

Our own products

Built for ourselves first

We build our own products with the same method we sell. They get the same tests, the same audits and the same refusal to state what nobody has confirmed.

Our own product

A booking and enquiry receptionist that knows your business end to end

Answers a small business's customers on its own WhatsApp number, from its own price list, and holds the slot in the database before it says "booked".

500+ appointments booked so far and 2,000+ enquiries answered in chat. In early access with a few clinics and car detailing shops. For Indian small businesses, priced monthly.

Ask for a demo

Client build

A voice agent that makes collections calls

Three agent personalities routed by risk score, promises extracted from the conversation into structured commitments, payment links sent during the call, and full transcript logging.

English today, shown as a live demo on sample data.

Ask for a demo

Lab

A strategy foundry

Paper trading only. The model writes strategies at design time; they run as deterministic code and must beat honest benchmarks net of costs.

Built in two days, with 129 tests and eighteen scanners. A 57-agent adversarial audit on 17 Jul 2026 raised 50 findings: 44 confirmed, 6 refuted, 15 fixed the same day. Deployment needs a single-use token and an operator's click.

Questions

Technical questions

What do the agents know before they start on our project?

Three things, read before the task: the expert's working memory, which we call the digital twin; your project's own rules file, which sits in the repository you receive; and the notes kept from earlier sessions on your project.

Do the agents work unattended?

For long stretches, yes, and by design: the answers they would otherwise ask for are already written down, routine actions do not wait for a click, and they check their own work with tests and a browser. On a recent monthly average, about 29 runs go half an hour or more with no person involved. They stop for what only a person should decide: putting anything live, spending money, deleting something that cannot be brought back, or stating a fact nobody has confirmed. Both lists are set out above.

Which AI do you build with?

Claude Code, Anthropic's coding agent, for the building. The models that run inside your system are a separate choice, made per project on measured results.

Is the chat on this site an AI?

No. It is a list of written answers, matched to your words in your own browser. No model reads your question, and a question with no answer goes to a person. It is shown at work above.

Who reviews what the agents write?

The expert, every change. Findings are acted on when reproduced, not when argued. What was removed is recorded.

How are agents stopped from inventing behaviour?

Design first, constraints in the database, tests written with the code, reproduction before acceptance, and a deploy that needs a human's click. Our strategy foundry enforces the last one with a single-use token so the invariant does not depend on the model obeying a prompt.

What happens when the model is wrong?

It is designed for. Deterministic paths answer first. When the answer is not in the approved material the system says so and routes to a person rather than guessing. Every answer carries provenance. On our booking product the measured booking path made zero model calls.

Which parts are deterministic and which are model calls?

Stated per feature in the design note and visible in the code. On our booking product there are four tiers, the model is tier four, and the ratio is measured.

Which models do you use, and are we locked in?

Whichever is right for the constraint, chosen on measured results, not brand. We build bring-your-own-model where we can. Licence terms are checked as part of the choice, not after it. That is about the models inside your system; for the building itself we use Claude Code.

What does a pull request look like?

A design note, the change with its tests in the same commit, migrations with the constraint in the database, and a line on what was refused and why. The full shape is set out above.

Can we run the test suite ourselves?

Yes. It is in the repository you receive, with the fixtures and the instructions to run it on Linux.

Can our own developers take it over?

Yes. The code is yours on payment, with its tests, documentation and runbooks, so your own developers can take over whenever you like. The hand-over also includes the build log and the backup and restore procedure.

Other questions: the main questions, and questions about the price.

Technical brief

Ask for the technical brief

The workflow, the test strategy, the security scope, hosting options and the hand-over list, in two pages.

Or ask us to show you, on a call

  • The rules file and the standing notes the agents read on one of our own products
  • A design note, and the change that was built from it
  • A test that failed before a fix and passes after it
  • The 108 answers that were read by hand, beside the grader's marks
  • The security review of our booking product, finding by finding