Data & integrations
When systems use the same word for different things
Separate identities, relationships and source meanings before automating the connection.
Two systems can use the same word and mean different things. A “location” may be a physical site in one system and a logical stock pool in another. Joining those records by name can make an integration look correct while changing the meaning of its data.
Before automating the connection, establish which objects the workflow needs to distinguish. Give them stable identities, record their relationships and keep external identifiers tied to their source.
The aim is a model that preserves the meanings needed for the task. It should be no more elaborate than that task requires.
Start with an event
Consider a constructed example: an order arrives through an online channel and is fulfilled from a depot. A provider describes both the channel's allocation pool and the depot as “locations”. The application needs to preserve several separate facts:
- where the order came from;
- which stock pool was used to allocate it;
- which physical site fulfilled it;
- which provider codes identify those objects.
One location field cannot reliably answer all four questions. A channel migration should not appear to move inventory, and a change of depot should not rewrite the source of demand.
Write down the events the software must support before choosing tables. An order received, stock reserved and a shipment dispatched involve different facts. The same object may participate in each event without serving the same role.
This makes the model easier to challenge. Ask what happens when an order splits across sites or a provider changes its codes. If the proposed design cannot describe the event clearly, it needs more work before data is imported.
Separate concepts that change independently
A physical site, a stock pool and a sales channel are useful distinctions in this example. They are not a universal schema that every product must adopt.
A site can perform several roles, such as storage, dispatch and returns. A stock pool may span sites or represent an allocation rule. A channel describes how demand arrives. Model the relationships the supported workflow needs without making the provider's terminology the application's permanent vocabulary.
Apply the same test elsewhere. An organisation can hold several commercial accounts. An address can change without creating a new organisation. A stock-keeping unit and a production batch may need separate identities because one describes the item and the other tracks a particular group of units.
Separate concepts when their rules, relationships or histories differ in ways the software must preserve. Avoid adding speculative layers merely because another system might eventually need them.
The interface can still provide a concise customer or product view. A clear internal model need not expose every distinction on every screen.
Keep identity separate from labels
Names change, repeat and contain errors. They help people recognise a record but often make poor identifiers.
Use a stable internal identifier where the domain needs one. Store external identifiers with their source system and relevant account or namespace. An identifier that is unique within one provider account may not be unique across all accounts.
A mapping should say which internal object an external record represents. Record validity periods or previous mappings when changes affect historical interpretation. For ambiguous matches, preserve the candidate and the reason for uncertainty rather than silently merging records.
Rules or models can propose a match. Similar spelling does not establish identity. A consequential merge needs checks appropriate to what it will change, and a review path when the available evidence cannot decide it.
In the order example, the provider's location code remains useful. It simply becomes an external reference mapped to a site or stock pool, rather than defining what “location” must mean everywhere.
Decide who owns each value
One system may own a description while another supplies shipment events. Assign responsibility at the level needed by the integration instead of declaring one platform authoritative for everything.
For a material value, identify the allowed writer, update direction, expected delay and response to conflict. Decide what the application shows when the source is unavailable or stale.
Timing matters. Two systems can disagree legitimately while a transfer is in progress. Reconciliation must account for that window instead of treating every difference as an error or ignoring differences indefinitely.
Keep source data available where it is needed to investigate a transformation. Record corrections and their basis when they affect reporting or traceability. The goal is to explain how a value reached its current state without retaining unnecessary copies of sensitive records.
Preserve history through change
A display-name correction differs from an organisation merger or replacement item. The model should represent the change that actually occurred.
Updating a label may be enough for a spelling error. A retired item may need a successor relationship. A duplicate record may point to a canonical record while old references remain traceable. Use effective dates where reports must reflect the structure that existed at the time of an event.
Do not add a complete event-sourcing system simply to preserve a few important changes. Choose the lightest history mechanism that answers the product's real questions and supports recovery from a mistaken correction.
Test agreement, not just successful requests
An API success response means a request completed according to that interface. It does not prove that both systems now represent the same business state.
Check meaningful relationships and totals. Active items may need provider mappings; pending events may need a maximum age; counts may need to agree after a defined processing delay. Make discrepancies understandable enough that someone can investigate them.
Test changes as well as current data: a new account, split fulfilment, a retired object, an uncertain duplicate and a renamed external identifier. These cases expose assumptions that a clean initial import can miss.
Building software alongside the people who use it explains how to uncover these meanings in practice. Follow the task, distinguish the objects it depends on, and keep that meaning stable as data moves between systems.