← All articles

AI & agents

Building software with agents

How understanding, clear decisions and short feedback loops help us work effectively with coding agents.

Coding agents can make implementation much faster. They can investigate a repository, propose changes, write code and run checks. Getting a useful result still depends on understanding the problem and recognising whether the result fits it.

We have worked with language models since GPT-3.5. As their capabilities have improved, our attention has shifted towards subtler failures. A change can run, pass its tests and look reasonable while resting on an assumption that nobody intended. The agent has filled a gap in the request with a plausible answer.

Clear communication matters at every stage. We need to explain the important relationships, identify decisions that remain open, and give the agent a way to check its work. We also need a quick way to inspect the result ourselves. Otherwise, time saved writing code can be spent untangling decisions or repeating slow rounds of review.

Understand what the system needs to represent

Before asking an agent to build, we need an account of the system that we can question. What are the important entities? How do they relate? Which rules must hold, and where does behaviour depend on the circumstances?

Agents can help trace these relationships and turn them into a structured description. The difficult part is deciding whether the description captures the work. An assumption about who owns a record, what a location represents or which role can make a change may shape the whole implementation.

In operational-software work, we needed to represent different entities, locations, roles and items. Understanding how people used the system revealed which distinctions the data model needed to preserve. Developing that foundation took sustained investigation, careful design and iteration before it could support broader use.

The challenge was deciding which distinctions mattered. Too many special cases would make the system difficult to extend. Too much abstraction would obscure meaningful differences. An agent can propose either kind of design fluently; judging it requires knowledge of the operation and examples that test its limits.

This does not require every detail to be settled upfront. It requires us to know which assumptions carry consequences, and which questions need investigation before dependent work proceeds. Working alongside the people who use the software is part of that investigation.

Make research part of the development process

Agents can quickly build visual explainers to help you understand how a system works and walk clients through it.

Start by establishing what actually happens. Trace a request through the interface, validation, storage and external services, including the return path and failures. Agents can help turn those findings into a map that makes dependencies and missing steps easier to inspect and discuss. Check it against the code and observed behaviour.

Parallel agents can investigate separate questions: one traces a data path, another examines tests, and another inspects an integration. Ask for evidence from the code and reconcile their findings. Several plausible summaries can disagree about the same system.

Research extends beyond the repository. Current documentation explains supported behaviour; users can explain an unusual rule. A small experiment may resolve uncertainty more cheaply than a large implementation.

Inspect the delivery process too. Which tests run locally and in continuous integration? What important behaviour is missing from them? Can the change be deployed and reversed? Coverage figures alone do not answer those questions.

For integrations, check what enters and leaves the system, where permissions are enforced and how failures are handled. Look for exposed credentials, unnecessary access and sensitive information in logs. Automated scans help, but cannot establish that the whole design is secure. Distinguish confirmed behaviour from inference and unresolved questions.

Explore the full atlas ↗, or get in touch to ask how it was made.

Plan around dependencies and evidence

A useful plan explains what we have, what we need and how we intend to get there. It also identifies choices worth comparing. A faster implementation might introduce a dependency that complicates maintenance. A flexible design might make permissions harder to reason about.

Compare robustness, efficiency, safety, security, privacy, delivery time and maintainability where they affect the decision. Keep the assessment proportionate. Some choices need investigation; others can be made quickly because they are inexpensive to reverse.

There are two ways to parallelise the work:

  • Divide one task between agents. Map its dependencies, assign independent parts and sequence the work that must wait. Agree shared interfaces and data meanings before implementation branches out.
  • Supervise agents on separate tasks. While one agent works, you turn to another task. This can use waiting time well, but you carry the cost of switching context and reviewing several streams of output.

Efficient parallel work depends on understanding what each change touches: files, interfaces, data and behaviour elsewhere in the system. Identify overlaps before assigning work. Agree who handles shared changes and which tasks must wait. For concurrent edits in one repository, use separate Git worktrees and focused commits. Worktrees isolate working files; check whether tasks also share databases, services or deployment environments. Merge in dependency order, resolve conflicts against the intended behaviour, then review the merged changes and test the combined paths. A clean merge can still combine incompatible assumptions; understanding the system is what lets you recognise them.

Agents can take time to investigate or implement a change. Each time you switch tasks, you need to recover the context and decide what needs attention. Learn how much parallel work you can comfortably supervise; that capacity varies with the tasks, your energy and the time of day.

Some work needs your full attention: watching outputs, inspecting changes and correcting the direction as it develops. A long series of repetitive edits or a straightforward implementation with clear checks can often run until an agreed checkpoint. While that runs, you might review completed code or think through the next steps over a snack or a glass of water.

Each stage needs a gate: the evidence required before dependent work proceeds. A data-model decision might need representative cases worked through. An integration might need a successful round trip and a demonstrated failure path. A user-facing change needs inspection in the interface where people will encounter it.

Define completion before implementation: expected behaviour, important exceptions, relevant checks and who will review the result. Revise the plan when new evidence changes it.

Test what needs to stay correct

Start with what a change could break and which behaviour people depend on. Tests should protect those requirements: calculations, permissions, data transitions and important user journeys. A smoke check that the application starts is useful, but does not establish that those behaviours work.

Agents can generate many tests that repeat the implementation's assumptions. Ask what failure each test would catch. Keep checks that detect meaningful regressions, and avoid tying them to incidental details that should be free to change. More tests also mean more maintenance.

Run fast programmatic checks first, then exercise integrated journeys. Scripted browser tests can repeat known interactions; computer-use agents can also explore the interface and report what happened. Check resulting records and downstream effects as well as visible success messages.

Choose when checks run as carefully as what they cover. CI runner minutes, deployment builds, database queries and paid API calls all have costs. Run fast checks on each change; trigger expensive integration tests when affected code, configuration or dependencies change, or at an appropriate release gate. Account for shared dependencies when selecting tests. Keep a scheduled or pre-release check of the complete journey to catch gaps in that selection. Each run should answer a specific question about whether the work can proceed.

A failing test needs interpretation. It may reveal a regression, a changed requirement, unreliable test data or a flawed check. Trace the cause before changing the code or the expected result. Do not let an agent redefine success simply to make a test pass.

Before client delivery, personally review the changed features and the important user journeys they affect. Exercise relevant loading, empty, error and permission states. Automated results support that review; they cannot replace your understanding of the system or your judgement of the experience.

Keep the visual feedback loop short

Visual work makes the limits of a written specification particularly apparent. A page can contain every requested element and still feel awkward. Moving between pages may shift the reading position. A transition may be technically smooth but poorly timed. Small inconsistencies in spacing and motion can make the whole interface feel rough.

In our experience, agents have reduced the effort of trying interface and animation changes. That makes a fast review loop especially valuable: we can explore more possibilities, provided we can see and judge them without a lengthy setup each time.

In one component-development task, we used Storybook to isolate a custom component and iterate on it quickly. Voice transcription let us narrate observations as we reviewed. We then translated those observations into precise requirements and investigated the code when a visual issue was difficult to explain.

Storybook stories describe components in particular states and configurations, making those states easier to revisit and test. That helps separate a component problem from the surrounding application. The component still needs review in context, where routing, layout, data and neighbouring elements affect the experience.

“Something jumps here” can be the start of useful feedback. Specify the action, the visible change and what should remain stable. Then investigate possible causes. A layout shift might involve font loading, container dimensions, scroll restoration or component state; changing an animation duration will not resolve every kind of movement.

Use a consistent visual language and shared components within the product. This reduces the number of independent design decisions and helps related screens behave coherently. Keep representative states easy to open locally, including loading, empty, error and long-content states.

Check the sizes and browser or device versions the product supports. Viewport controls help inspect different dimensions, but a resized preview does not reproduce every device's behaviour. Review keyboard interaction and reduced-motion settings too. Automated checks can identify regressions; a person still needs to assess continuity, emphasis and feel.

UX is part of the system explores how those individual interactions fit the wider product.

Develop better judgement through iteration

Working effectively with agents involves learning where they are likely to make assumptions. In our experience, this includes proposing more schema machinery than the task needs, missing small visual inconsistencies, or committing to a familiar solution before exploring alternatives.

That experience informs how we ask questions and review answers. Ask what a proposed abstraction makes possible and whether the requirement needs it. Ask which assumption a test exercises. When a result feels wrong, identify the observable difference before requesting another broad rewrite.

Technical grounding helps us investigate a result and judge the alternatives. Imagination and taste matter when several solutions work but produce different experiences. Clear communication gives the agent enough direction to develop the one we intend.

A useful place to start is one bounded change. Establish the current behaviour, name the unresolved decisions and decide how to inspect the result. Let the agent work against those requirements, review what happens, and adjust. Judge the efficiency of the whole loop, including investigation, correction and verification.