When a legacy app is insecure at the architecture level, patching it in place can cost more and carry more risk than replacing it, and every week it stays live keeps the exposure open. The safer path is to build the replacement beside it, move users in measured steps, and never touch the old system’s data on the way out. This is how Shovel Solutions is retiring its legacy hauling marketplace, and what the move has taught us so far.
Shovel Solutions runs a regional marketplace for dirt and aggregate hauling: job creators pay for delivered loads, and pits, independent truckers and fleets get paid. Our assessment of the legacy platform found that its weaknesses came from one design choice, a real-time, trigger-driven document store with no transactions and no idempotency. You cannot add correctness to that one function at a time. The decision was to retire it.
The migration is in progress. The new platform is deployed to production but not yet live, and the steps below are the strategy as built and rehearsed so far, not a finished result.
Where the migration stands: the first two steps are running; the rest have not happened yet.
Build the destination on correctness first
A migration is only as safe as the system it lands in, so the new platform’s foundations came before any user moved.
- A double-entry ledger in PostgreSQL. Every money movement is a journal whose debits and credits must balance, and entries are appended, never edited. Balances are derived from the ledger rather than stored in a field that can drift. This replaces the mutable balance the legacy system kept.
- Idempotency on every external call. The design requires an idempotency key on every payment-provider request and de-duplicates incoming webhooks by event ID. One example of why the key’s scope matters: creating a customer at the payment provider once had no key, so a retry after a failed local write created a second customer. The key is now scoped to the account rather than the email address, because addresses come from the caller and two accounts can share one.
- Restores that are rehearsed, not assumed. A production point-in-time restore has been rehearsed and took about 8 minutes. A drill runs weekly: it reads the actual restorable window from the server, restores to a new server, compares the two, and then deletes the restored copy behind a name guard, so a failed drill cannot leave a server running up charges. Extending the drills to every database and to a cross-region restore is the next step.
-- Illustrative: a journal posts only if it balances.
CREATE TABLE ledger_entries (
id bigserial PRIMARY KEY,
journal_id uuid NOT NULL,
account_id uuid NOT NULL,
amount numeric(14,2) NOT NULL, -- debits positive, credits negative
created_at timestamptz NOT NULL DEFAULT now()
);
CREATE FUNCTION assert_journal_balanced() RETURNS trigger AS $$
BEGIN
IF (SELECT sum(amount) FROM ledger_entries WHERE journal_id = NEW.journal_id) <> 0 THEN
RAISE EXCEPTION 'journal % does not balance', NEW.journal_id;
END IF;
RETURN NULL;
END $$ LANGUAGE plpgsql;
CREATE CONSTRAINT TRIGGER journal_balanced
AFTER INSERT ON ledger_entries
DEFERRABLE INITIALLY DEFERRED
FOR EACH ROW EXECUTE FUNCTION assert_journal_balanced();
Because the trigger is deferred, it runs at commit, so a journal’s entries can be inserted one at a time inside a transaction and are checked together.
The legacy database is read-only to us
The migration’s firmest rule is that the legacy database is never written to or deleted from. The old app keeps serving its users, unchanged, until they move.
That rule makes the migration reversible. If something on the new side is wrong, the source of truth is still intact, and nobody has to reconstruct the old state from a half-run script. It also keeps the legacy system out of scope for new changes, which matters when the reason for leaving is that changes there are risky. A “purge and re-migrate” step was considered for the go-live sequence and ruled out.
A live switch, with delta loads behind it
Production runs with a live switch off and a legacy import on. Data is copied across in bulk first, then topped up with delta loads that bring over only what changed since the last run. Going live is an ordered act: run a final delta load, turn the import off, turn the switch on.
Delta loads are where the subtle bugs live, because the new system starts making its own decisions before the old one stops. We found two of the same kind before the final delta load:
- A delta load was rolling back claims a crew had already completed in the new system.
- The same pattern was re-activating memberships that closing an account had deliberately revoked. A closed account would quietly come back on the next sync.
The second was found by searching for the first one’s pattern elsewhere. The general rule: once a record has state the new system owns, a delta load may add to it but must never overwrite it. Each import rule should be checked against that question and tested with a sync that runs after a user has acted.
Pace the invitations to the slowest limit
Moving users means telling them, and an invitation that fails silently is worse than one that never went out. The new platform runs its own identity provider, which accepts five password-recovery requests an hour per source address. That limit protects real users from abuse. An invitation run comes from one machine, so it shares that one counter.
The original plan assumed a bulk invitation would finish in an afternoon. At five an hour, roughly a thousand users takes more than a week. Worse, an unpaced run would have delivered a handful of emails and reported every one of them as accepted, because “accepted” is the only answer the service gives.
The fixes were plain:
- The run is paced to the limit, and a run in which anyone failed reports a failure.
- Each run reads how much of the hour’s allowance is already spent before it sends anything. Without that, a second run inside the same hour could record people as invited who were never mailed, and no later run would pick them up again.
- Every run ends by listing who it did not reach.
The plan has since changed in a way that makes this easier. Existing accounts will no longer move in one bulk invitation. They will come across as needed, and people will be pointed to the identity provider’s own forgotten-password page. The pacing and failure reporting stay, so whatever is sent reports its results accurately.
What this approach buys you
- No big-bang date. Users move in steps, and any step can stop without stranding anyone.
- A reversible path. The legacy system is untouched until it is decommissioned, after cutover.
- Correctness you can test. The ledger, idempotency and restores are each checked by something that runs, not by a design document.
How we can help
We assess legacy apps, decide with you whether to patch or retire them, and run the migration so users move in measured steps without a risky cutover. The work starts with the destination’s correctness and a rehearsed restore. See our app stabilization service.