Microservices · FinTech · Distributed Systems
Designing Microservices for Financial Systems
Service boundaries in a payments platform are not an architectural preference, they are a correctness decision. How money, state and failure change the usual microservices advice.
Most microservices writing is about deployment independence and team autonomy. Both matter. Neither is the reason you split a financial system.
In a payments platform, the boundaries you draw determine where a transaction can be observed, where it can be retried, and where it can silently go wrong. That makes service decomposition a correctness concern before it is an organisational one, and it changes several pieces of the standard advice.
Start from the transaction, not the domain#
The usual decomposition exercise starts with domain modelling: find the bounded contexts, one service each. It is good advice and it under-specifies the financial case, because in a financial system the question that decides the boundary is narrower:
Which operations must succeed or fail together, and which must be allowed to resolve independently?
Anything inside the first group wants to be inside one service and, ideally, one database transaction. Anything in the second group is a candidate boundary. Split across the first group and you have replaced a transaction with a distributed saga, which is sometimes correct and always more expensive than people expect.
A practical heuristic: if explaining the failure recovery for a proposed boundary takes more than a paragraph, the boundary is in the wrong place.
Three outcomes, not two#
Ordinary services have two outcomes. Financial services have three, and the third one is where the engineering is:
| Outcome | Meaning | What the system must do |
|---|---|---|
| Success | Value moved, confirmed | Record, notify, close |
| Failure | Value did not move, confirmed | Record, release any hold, close |
| Indeterminate | Unknown, timed out, no callback, partial | Keep the record open and actively resolve it |
The indeterminate state is not an edge case. Network timeouts, missing callbacks and provider-side successes that never reach you are routine at any real volume. A service that models only success and failure will eventually record one of them incorrectly, and a wrong record in a financial system is worse than no record.
This has a direct architectural consequence: a transaction's lifecycle outlives the request that started it. Something has to own it afterwards. A status poller, a callback handler, a reconciliation job. That owner is usually its own service, and forgetting to plan for it is the most common way a payments architecture ends up incomplete.
Idempotency belongs at the boundary, not in the client#
Every service that mutates financial state should accept a client-supplied idempotency key and store what that key produced. Not "should consider", should.
POST /v1/transfers
Idempotency-Key: 6f1c0b2a-8e3d-4f7a-9d21-0c4b9e7f1a55
Content-Type: application/json
{ "amount": 250000, "currency": "NGN", "destination": "..." }
The rule is that a repeat of the same key returns the original result. The same transaction identifier, the same status, rather than performing the operation again. The subtlety worth getting right: a key that arrives while the first request is still in flight should not start a second operation either. That means the key is reserved on receipt, not on completion.
Pushing this responsibility to callers does not work. Callers retry for reasons outside their control, proxies, client libraries, users. The boundary has to be safe on its own.
Each external dependency gets an adapter#
In a system that integrates banks, mobile-money operators or billers, the single highest-leverage structural decision is to isolate every provider behind its own adapter with its own mapping, timeout profile and error translation.
The benefits compound:
- Onboarding is contained. Adding a ninth provider requires no reasoning about the first eight.
- Failure is contained. One provider degrading does not become a platform incident, because its timeouts and retries are its own.
- Semantics are explicit. Each adapter declares, in code, whether its provider's success response means settled or accepted. That distinction gets lost when integrations are inlined into business logic.
The adapter layer is also where you put per-provider circuit breaking. A provider that is failing fast should be failing fast for everyone at once, not rediscovered independently by every caller.
Reporting is a separate read path#
The operational question, did this customer's payment arrive?, and the analytical question, how did this provider behave last week?, have different latency requirements and different access patterns. Serving both from the same store and the same service means the analytical query eventually degrades the operational one, usually at the worst moment.
Separating the reporting path is not premature optimisation in this domain. It is the acknowledgement that someone will run an unbounded query against your production data eventually, and it is better that they do it somewhere that cannot hurt a transfer.
What does not change#
Two pieces of conventional advice survive contact with financial systems intact:
Do not start with microservices. A boundary drawn before you understand the transaction semantics is a boundary you will move, and moving a service boundary in a system that holds money is painful in a way that moving a module is not.
Prefer a boring synchronous call to a clever asynchronous one until the asynchronous version is justified. Every queue you add is another place a transaction can be sitting when someone asks where it is.
The theme underneath all of this: in financial systems, architecture is largely the practice of deciding where the truth lives and how you would prove it. Get that right and deployment independence follows. Get it wrong and no amount of deployment independence helps.
Next article
Building Reliable Payment and Transfer Integrations
A field guide to integrating banks, mobile-money operators and billers, written from the perspective that every provider will eventually behave in a way its documentation does not describe.
Contact
Let’s Build Something Meaningful
I’m open to senior and lead engineering roles internationally, backend, full-stack, architecture and platform work. The fastest route is email; LinkedIn works just as well.