Product engineering
Staying with the build
Make support, recovery and knowledge transfer practical parts of looking after software.
Once people use a system, engineering work changes. Questions come from real tasks. Integrations fail in ways the prototype never encountered. A new member of staff needs help, a dependency changes, or an apparently small request exposes a gap in the design.
Supporting that software takes more than keeping its hosting bill paid. Someone needs to understand what people rely on, investigate problems and make changes without losing the parts that already work.
The original developer can provide that continuity, or help another team take over. Either arrangement should be deliberate and practical.
Decide who will look after the system
Discuss ownership before the build ends. The answer affects infrastructure, documentation and how releases are organised.
An ongoing engagement can combine development and support. A handover can transfer responsibility to an internal team. A transition period can give that team time to practise with help available. Some systems need a split, with different people maintaining the application and a specialist component.
Make the boundaries understandable. Who receives a user report? Who can inspect the relevant logs? Who decides whether a request is a defect, a feature or an upstream data problem? Who can approve and release a change?
Responsibilities without access are difficult to exercise. Confirm that the people maintaining the system can reach its repositories, service accounts and recovery tools through appropriate permissions. Record where credentials are managed rather than putting secrets in a runbook.
Learn from the support work
A user report usually begins with an effect: a number looks wrong, a task is stuck, or something no longer appears where expected. Reproduce the task and trace the issue before deciding which component is responsible.
The cause may be code, source data, an integration or an unclear interface. It may also be a legitimate change in how the organisation works. A technically correct system can still require a product change.
Keep recurring manual repairs visible. If someone regularly reconciles records or restarts a job, record the pattern and its cost. It may justify better validation, a recovery control or a change to the source process. Do not let repeated intervention become an invisible part of normal operation.
Onboarding is another useful source of feedback. A new user may reveal assumptions that experienced staff no longer notice. Improve the interface or guidance where possible instead of relying indefinitely on an explanation from its developer.
Make failures understandable and recoverable
Monitor the conditions that matter to the task. Service availability and errors are a start; integrations may also need checks for stale data, missing events and unresolved work.
An alert needs an owner and a useful response. More notifications do not help if nobody can distinguish a routine delay from a failure that blocks users. Google's SRE monitoring guidance explains the value of monitoring with clear purpose and actionable signals.
Recovery should include the work affected by the failure. Bringing a service back online may leave partly completed records or uncertain actions behind. Establish how to find and reconcile them, and how a retry avoids duplicating an effect.
Practise important procedures in a controlled environment. Restore a backup, exercise a rollback, or have another engineer follow the release instructions. The exercise often reveals missing permissions or assumptions that a document review misses.
Collect enough diagnostic information to explain events, with suitable access and retention. Do not copy whole user requests into logs merely because they might be useful one day.
Keep changes small enough to check
Dependencies and providers continue to change whether or not the application receives new features. Maintain a practical inventory of important services, versions and expiry dates, with someone responsible for reviewing them.
For a material change, identify what could be affected, test it in the appropriate environment and define a recovery path. Inspect the result after release. Urgent fixes may compress the process, but they still need a record of what changed and follow-up on any checks deferred.
AI components need task evaluation as well as software tests. A model, prompt or retrieval update can alter output quality, tool selection and cost. Compare representative cases before treating an upgrade as an improvement.
Keep documentation with the change it explains. A new integration should bring its failure behaviour and ownership with it. A deployment change should update the release instructions. This is usually easier than reconstructing everything before a handover.
Hand over through practical work
A repository and a meeting do not show that another team can operate a system. Plan the transfer around tasks they will need to perform.
Have the receiving team run the application, make a small change and release it through the normal process. Ask them to investigate a prepared failure and exercise an appropriate recovery procedure. Use their questions to repair both the documentation and any avoidable complexity in the system.
Transfer the reasons behind important decisions too. Explain unusual data rules, provider constraints and known limitations. Another engineer should be able to distinguish a deliberate boundary from accidental implementation detail.
Confirm account ownership, access and escalation paths at the agreed point. Remove delivery-team access when it is no longer needed. A transition can remain available for a defined period without leaving ownership ambiguous.
Agree a support scope people can rely on
State which systems, activities and hours the arrangement covers. Separate acknowledgement and investigation targets from a promise to resolve a problem: recovery may depend on a provider or on repairing data.
Give planned improvements and urgent incidents a way to compete for capacity. An unlimited stream of features, questions and repairs is hard for either side to plan around.
Review the arrangement as the product changes. Continuing support should leave the software easier to understand and maintain. A handover should leave the receiving team able to act. Both can provide continuity when the responsibilities are matched by knowledge, access and time.