Skip to content

Designing API integrations for the day the other system fails

The happy path is not the integration. The integration is what happens when the other system is slow, down or wrong.

Written by Tarun Mookhey · CTO / Principal Engineer ·

Most integrations are demonstrated on the happy path: call the API, get a response, move on. In production, the other system will eventually be slow, unavailable, rate-limited or simply wrong. An integration is not finished until it has a defined behaviour for each of those days.

Decide who owns each piece of data

For every object that moves between systems, decide which one is the source of truth, and in which direction changes flow. Two systems that both believe they own a customer record will disagree sooner or later. Writing the ownership down before building anything removes a whole category of later argument.

Authentication is part of the design

Integrations act as someone: a service account, or the user. That choice affects permissions, audit trails and what happens when a token expires or is revoked. OAuth2-connected user accounts, for example, need a plan for refresh failures and for users who disconnect.

Webhooks and polling

Webhooks are timely but unreliable on their own: deliveries can be missed, delayed, duplicated or arrive out of order. Polling is slower but easy to reason about. Many robust integrations use both: webhooks for speed, a periodic poll to catch what was missed.

Retries need idempotency

Retrying a failed call is only safe if repeating it cannot do harm. Give each operation a stable identifier so that sending it twice has the same effect as sending it once, and back off between attempts so a struggling system is not made worse. Decide when to stop retrying and where the failed item goes instead.

Expect duplicates and partial failure

Assume the same event can arrive more than once, and that a multi-step operation can succeed in the middle and fail at the end. Design each step so it can be repeated, and record enough state to resume.

Plan reconciliation

However careful the design, systems drift. A scheduled reconciliation that compares the two sides and reports (or repairs) differences turns silent corruption into a visible, fixable item.

Respect limits, and make failure visible

Rate limits, quotas and timeouts are part of the contract. Build in throttling, and make failures observable: structured logging, alerting, and a way for someone to see what is stuck and retry it by hand. A manual recovery path is not a sign of weakness; it is what makes the automated path trustworthy.

A checklist

  • Source of truth defined for each object
  • Authentication and token failure handled
  • Webhooks backed by polling or reconciliation
  • Idempotent operations with back-off
  • A place for items that keep failing
  • Monitoring and a manual recovery path

Related: API & systems integration · TickMessage case study · Ecommerce and fulfilment case study

Have software that needs building, connecting or untangling?

Describe the problem as it exists today. A finished specification is not required.

You will speak directly with the senior engineer responsible for assessing the work.

Discuss a project