API Design · Backend Engineering · Architecture
Lessons From Building Production APIs
An API is a promise you have to keep after you have forgotten making it. What ten years of building integration surfaces teaches about versioning, errors, contracts and the clients you cannot upgrade.
The difference between an API and a production API is that somebody else's roadmap now depends on yours.
Everything below follows from that. These are the lessons that have held up across payments middleware, logistics platforms and property services, the places where an interface was not a convenience for my own front end but a contract other teams, partners or mobile clients built against.
Your oldest client defines your constraints#
Web clients update when you deploy. Mobile clients update when a user decides to, if ever. Partner integrations update when the partner schedules the work, which may be never.
This single asymmetry drives most good API decisions:
- Adding is safe. Removing and renaming are not. New optional fields are free. Removing a field, changing its type, or tightening validation are breaking changes even when nothing in your test suite notices.
- Defaults are permanent. A default value you chose casually becomes the behaviour thousands of unchanged clients depend on.
- Silence is a contract. If your endpoint currently ignores an unknown field, somebody is sending one.
Practical test before any change: if a client written two years ago called this endpoint today, what would happen? If the answer is "it breaks," it is a new version, not a fix.
Design the error taxonomy first#
Success responses design themselves. Errors are where APIs become unusable, and they are almost always designed last.
A usable error tells the caller three things: what category of problem this is, whether retrying could help, and what specifically went wrong.
{
"error": {
"type": "validation_error",
"code": "amount_below_minimum",
"message": "Amount must be at least 100.00 NGN.",
"retryable": false,
"field": "amount",
"reference": "req_8f2c1d9a"
}
}
Four things make this work:
type is a small closed set. Callers switch on it. Keep it to a handful,
validation, authentication, authorisation, not found, conflict, rate limit,
provider error, internal, and never add one casually.
code is stable and specific. This is the string that ends up in someone's
if statement. Renaming it is a breaking change; treat it like a public symbol.
retryable is explicit. Do not make callers infer retry safety from an HTTP
status. They will infer it wrong, and they will infer it wrong in the direction
that duplicates work.
reference is the thread back to your logs. When a partner reports a
problem, the first question is always "do you have the reference?" An API that
does not return one converts every support conversation into an archaeology
project.
Version at the point of breakage, not on a schedule#
Two failure modes, equally common. Some teams version everything continuously and end up maintaining six near-identical surfaces. Others refuse to version and break clients quietly.
The middle path: stay on one version until a change is genuinely breaking, then version deliberately and support the old one on a stated timeline. Most evolution (new fields, new endpoints, new optional parameters) needs no version bump at all if the design was additive to begin with.
When you do break, the deprecation needs three things clients can act on: a date,
a migration path, and a way for them to find out. A Deprecation header and a
changelog beat an email that reached the wrong person.
Make the common case obvious and the dangerous case explicit#
Good API ergonomics are mostly about where you put the friction.
- The common operation should require the fewest parameters and have sane defaults.
- The irreversible operation should require something extra, an explicit confirmation field, a narrower permission, an idempotency key.
If deleting is as easy as reading, the API is optimised for the wrong outcome.
Pagination is not optional, even when the list is small#
Every unpaginated list endpoint is a future incident. The data is small until it is not, and by the time it is not, clients depend on receiving everything.
Cursor-based pagination avoids the consistency problems of offset-based pagination under concurrent writes, and it degrades more gracefully at depth. Ship it from the first version even when the collection has four items.
Write the contract down in a form that executes#
An OpenAPI description that is generated from, or validated against. The implementation is worth substantially more than one maintained by hand, because the hand-maintained one drifts and nobody discovers it until a partner does.
The version of this that matters most: contract tests that run in CI. Not "does the endpoint work," but "does the endpoint still match the shape we published." That is the test that catches the accidental breaking change before it reaches a client who cannot upgrade.
Observability is part of the interface#
Callers cannot debug what they cannot see, and the support burden of an opaque API lands on your team.
The minimum useful set:
- A request reference returned on every response, success and failure alike.
- Status transitions recorded with timestamps for anything asynchronous.
- Rate-limit headers that tell callers their remaining budget before they exhaust it.
- A status page or health endpoint that is honest about partial degradation.
A partner who can answer their own question does not open a ticket.
The lesson under the lessons#
APIs fail slowly. A poorly-designed internal function is a bad afternoon; a poorly-designed public interface is a constraint on every subsequent decision for years, because the cost of changing it is borne by people who did not choose it.
Which argues for one habit above all the rest: before you ship an interface, write the integration guide for it. If explaining it honestly is difficult, the design is not finished. That document costs an hour and has saved me more rework than any other practice in this list.
Next article
Designing Microservices for Financial Systems
Service boundaries in a payments platform are not an architectural preference, they are a correctness decision. How money, state and failure change the usual microservices advice.
Contact
Let’s Build Something Meaningful
I’m open to senior and lead engineering roles internationally, backend, full-stack, architecture and platform work. The fastest route is email; LinkedIn works just as well.