← Back to work

Design note · Distributed workflows

Approvals between two services

How I designed the approval flow between two services at TÜBİTAK BİLGEM so that an approval can only apply to the data the approver actually saw, and so that a rejected action leaves nothing behind.

The setup

Two services handle the same travel expense claims. The owning service holds each claim and its version; accounting officers edit claims there. Supervisors work in a second service, which keeps a mirrored copy of the claim and sends their actions to the owner over Kafka. The owning service is the single source of truth.

What goes wrong without a rule

Some fields shouldn’t become final on the supervisor’s side until the owner accepts the action: an approval date, attached documents, a status change, a summary that goes on to another system. Without a clear rule for when data becomes final:

  1. The requesting service saves a field, then the owner rejects the action.
  2. Fields prepared before the decision get stored as if they were final.
  3. A document is generated for an action that is never accepted.
  4. An action is made against an old version and nobody notices until later.
  5. A failed action leaves its side effects behind.

The worst of these is the ABA problem. A supervisor opens a claim for ₺5,000 and leaves the screen open. The officer pulls the claim back, changes it to ₺50,000 and sends it to the supervisor again. The supervisor, still looking at the old screen, clicks approve. A check for “is this claim awaiting approval?” passes. The content is different.

The decision

Commands for actions that cross the service boundary, synchronous writes inside the owner, and optimistic versions carried end to end.

The version’s path:

  1. The supervisor’s screen reads the current version.
  2. The request sends it back with the action.
  3. The requesting service compares it with its mirrored copy.
  4. The same version goes into the Kafka command.
  5. The owner validates it against the real current version.
  6. The result carries the latest version back.
  7. The requesting service updates its mirror from the result.
  1. Supervisor’s screen: Opens the claimv1 · ₺5,000
  2. Owning service: Officer edits and resubmitsv4 · ₺50,000
  3. Screen to requesting service: approve, v1
  4. Requesting service to owning service: command, expects v1 Kafka
  5. Owning service: v1 ≠ v4nothing applied
  6. Owning service to requesting service: result: failed Kafka
  7. Requesting service to screen: claim has changed
The approval from the old screen is rejected and nothing is applied.

A result is more than a status

Success and failure are published by separate methods. A success can carry a changes envelope: everything the requesting side should now make permanent, such as the approval date or the final set of documents. A failure carries an empty envelope by construction, so the requesting side can’t persist anything for a rejected action, even by mistake.

The envelope is shared by every operation. A new approval step adds its own sub-object instead of changing the message shape, so result handling stays the same everywhere. A success can describe the complete new set of documents rather than a diff, so the requesting side replaces its set instead of patching it.

Prepared is not final

Some side effects have to exist before the decision, because the owner needs them to decide: a generated document, an uploaded file, a computed date. Producing them doesn’t make them final. If the action is rejected, whoever owns a side effect cleans it up; here that is the owning service, since it made the decision. A database rollback doesn’t delete an uploaded file, so the cleanup is explicit, and if it fails, the failure stays visible instead of being swallowed.

Edge cases

A stale screen

Rejected by the version check. The supervisor sees that the claim has changed.

A Kafka rebalance

A consumer that takes too long can lose its partition mid-message, and another consumer gets the same command. The first write has already bumped the version, so the duplicate fails with an optimistic-lock exception instead of applying twice.

Two people at the same moment

The officer’s synchronous write can reach the database before the supervisor’s command even if the supervisor clicked first. Nothing is corrupted; the late side gets a “data has changed” message.

Hearing your own messages

One topic carries messages in both directions. Every message names the service that sent it, and each consumer skips its own.

A result for a deleted record

A result can arrive after its record was removed. The consumer logs it and moves on, instead of failing on every retry for a record that will never exist again.

Alternatives we rejected

Everything through Kafka

Over-engineering for the owner’s own writes. And without a version from the screen, putting commands in order still doesn’t fix the ABA case.

A synchronous REST check before sending the command

It couples the services: an outage in the owner breaks the other side’s UI. The record can also change between the check and the write, so the race is still there.

Checking status only

Misses the ABA case entirely.

A distributed lock

A Redis mutex or similar would add infrastructure and failure modes to get a guarantee the existing optimistic locking already provides.

What it costs

Supervisors don’t get an immediate answer. The UI has to say the action was received and report the outcome when the result arrives. Under contention the synchronous side wins, which is acceptable because the losing side is told, not overwritten.

Adding a new approval step

I also wrote an implementation guide so other developers reuse the pattern instead of reinventing it. It starts with one question: would it cause a problem if the requesting service wrote this change immediately? If the answer is yes, the step follows this flow:

  1. Identify the root record and where its version comes from.
  2. Carry that version in the user’s request and in the command.
  3. Separate the action’s content from the fields that become final only on success.
  4. Add a sub-object to the result envelope for those fields.
  5. Decide who cleans up each side effect when the action fails.
  6. Test the version mismatch, the failure result, self-published messages and missing records.

More of my work · salihkaragollu@gmail.com