How we gave our AI agents a way to check their own work
The year is 2026, and AI has made its way into almost every corner of software development, with results ranging from impressive to embarrassing. One direction the industry finds especially promising is autonomous agentic development, where an agent takes a task written in natural language and writes the code on its own, start to finish.
There’s a catch. An agent that can’t verify its work is just guessing with confidence (even though, credit where credit is due, the guessing becomes better every month). It needs a way to know whether the task is actually done right, and that means tests.
When we, the development team, started moving CleanTalk’s development toward agent loops, it quickly became clear that our modest set of unit tests wouldn’t cut it. We needed something that exercises the real system, end to end.
Long story short: here’s how we built an E2E test harness for our API, and how it became part of our everyday development.
Project context
The project is a set of API methods. Other parts of our infrastructure call them, and so can regular CleanTalk users.
Some methods only read data and return it in a convenient shape. Others actively change the database, and some also reach out to other systems: KeyDB, payment gateways, and so on.
That mix makes testing tricky. For a method that writes data, a correct JSON response proves very little. What matters is what actually happened in the database.
Why E2E tests
An agent can run unit tests too. But a green unit suite proves that the pieces work in isolation, not that the system works as a whole. A method can pass every unit test and still fail on a real query because of some configuration or architecture issues.
Our rule of thumb: the higher the level of the check that runs after the agent’s work, the more we can trust the result. End-to-end tests sit at the top. They send a real HTTP request through the real method code into a real database, then check what actually changed.
That makes them the natural feedback signal for agent loops. The agent writes code, runs the suite, sees what broke, and fixes it without a human in every iteration, until all the tests are green.
The rest of this article is about what it took to make that possible.
Step 1: Keep the schema up to date
We have several MariaDB databases for different needs, and their structure changes over time. Schema changes can be initiated by different teams. So the first question was: how do we make sure the tests run against the schema that production actually has?
Our answer is a small utility that pulls the schema straight from the databases:
- It connects to each database with a dedicated, restricted database user.
- It collects every
CREATE TABLEstatement, plus functions and triggers. - It writes all of that into a
.sqlfile.
The result is a “recipe” for our databases: everything needed to recreate their structure, with no data. The harness builds its test databases from this recipe, as described below. When someone changes a table, we can re-run the utility and the tests would pick up the new structure. Doing this once a week is usually just fine.

Output of our schema dump tool.
Step 2: Design the harness
Since we only test API methods, each test needs to check two things:
- The response. Status, payload, error codes.
- The database state after the call. For example, if a method adds a new service to a user, the test confirms that the service record actually appeared in the database.
No UI is involved, so we don’t need a headless browser.
Testing is always a trade-off between isolation and efficiency. We want tests and suites to be unable to affect each other’s results. But isolation has a price: every fresh database or server costs time and resources, and too much of it makes the suite too slow to run after every change. Here’s the compromise we settled on.
Suites. All tests are grouped into suites. Each suite covers one API method and contains several tests. Every suite defines:
- the set of tables it needs;
- a set of starting data (fixtures).

One of our suite.php configuration files.
Fresh databases per suite. At the start of a suite, the harness creates new test databases with CREATE DATABASE. It then takes the suite’s table list and creates only those tables from the schema recipe. A suite usually needs only a handful of tables, so there is no point running hundreds of CREATE TABLE statements that no test will use. At the end of the suite, the test databases are dropped.
Reset before every test. Before each test, the harness puts the database back to its starting point: it clears the table data and loads the suite’s fixtures again. Every test starts from the same known database state.
Step 3: Run it
Here is what happens when a suite runs:

The simplified scheme of one test suite run.
For each suite, the harness starts a separate PHP server on its own port. Its database config points to the freshly created test databases.
The tests themselves are written with Playwright, used purely as an HTTP client and test runner. Each test resets the database to the fixtures, then sends a real HTTP request to that server. From there, the real code of the API method runs, exactly as it would in production. Finally, the test checks the response and the state of the database.
When the suite finishes, the harness stops the PHP server and drops the test databases. The next suite creates brand-new ones, and so on.
You can run a single suite or filter tests by name (more precisely, by their description). That keeps the feedback loop short when you’re working on one method.

Test run log.
Put the context in the repo, not in the prompt
An agent is only as good as the context it gets. Most of that context is the same for every API method, so instead of repeating it in every prompt, we keep it in the repository:
- README.md describes (among other things) everything above: the schema recipe, suites, fixtures, per-suite databases and servers.
- AGENTS.md holds our code style, our recommendations on which tools and patterns to use, and a guide to the Composer commands that run the tests.
The agent reads both before it starts. From there it can find the method’s code, the tables it touches, and existing suites to copy from on its own. What’s left for the prompt is one line:

Screenshot of a prompt for a test suite that was integrated later that day.
And here is what comes back:

A request like this usually produces a minimal suite that covers the most important scenarios. Each test carries a plain-text description of the scenario it checks.

Some of our tests.
Scenarios come with the tests
We don’t always write a list of scenarios up front. The test descriptions are the scenario list, and they arrive with working tests attached. Reading the test suite can give a descriptive idea of how the method should behave. A reviewer reads the descriptions and checks them against how the method is supposed to behave, not just against what the code currently does.
Then we often ask the agent a follow-up: which branches of the method’s behavior are still not covered by the test suite? If any of them matter, one more prompt closes the gap.
Why are we happy now
We are actively moving toward agent-driven development, and E2E tests are what let an agent work with far more autonomy. Instead of waiting for a human to confirm each step, it can check itself:
- Fewer handoffs. We don’t just assign a task; we give the agent a way to verify it did the task correctly.
- Legacy code stops being scary. When an agent (or a human) touches old code, regressions show up in the test run instead of production.
- Tests can come first. For new features, we write the tests, then let the agent implement the logic until they pass.
The result: work on our API methods has become 3–5 times faster, and we now spend our time on contracts and business logic instead of type-casting nuances. A method’s entire contract can be written down as a set of tests, and the code itself can be delegated to an agent.
The harness isn’t finished yet. Not every part of our system is wired into it, and some dependencies are still mocked. We’re working on it.
Wrapping up
The harness itself is not that complicated: a schema dump, a few CREATE DATABASE calls, a PHP server per suite, and Playwright sending HTTP requests. But together with a clear README and AGENTS.md, it turned “write tests for this method” into a one-prompt task, and made changing our API methods much faster and easier.
And the part we still find a little surprising: the first version of this harness took one day to build, and we built it with an agent too.
If you’re moving toward agentic development, a solid test harness can speed you up dramatically. Don’t put it off for too long: the sooner an agent can check its own work, the sooner you can trust it with more.
Leave a Reply