Skip to content
September 29, 2026· 2 min read

Designing a Database Migration You Can Resume

DB Orchestrator title card: resumable on-prem SQL Server to Azure SQL migration pipeline.

Moving one production database from an on-prem SQL Server VM to Azure SQL Database is a checklist. Moving many of them, safely, repeatedly, and without a person babysitting each one at midnight, is a system design problem.

This is the pattern behind an internal tool I built for exactly that. I’ll keep the specifics general, but the design ideas travel well.

The problem with a script

The obvious first version is one long script: stop writes, back up, restore, export, import, create logins, verify. It works right up until step six fails at 2 a.m.

Then you’re asking questions no script can answer:

  • Which steps already ran?
  • Is the source still accepting writes?
  • Can I rerun from here, or do I start over and hope?

A long migration will fail somewhere, eventually. The design has to assume that.

A state machine, stored in a database

Instead of a script, every migration is a job with an explicit state, persisted in a database:

Created → StoppingWrites → BackingUpSource → RestoringToMiddleman → PrepScripts → ExportingBacpac → Importing → CreatingLoginsAndUsers → Verifying → ScalingDown → ReadyForCutover → Completed

A few properties fall out of that:

  • Any step can fail into AwaitingMitigation. Someone fixes the cause, and the job resumes from the step that failed, not from the beginning.
  • Every transition is recorded, so the job’s timeline is an audit trail for free.
  • Cutover is always manual. The machine gets you to ReadyForCutover; a person flips the switch.

Don’t start until it’s safe

The riskiest moment is the start: migrating a database that’s still being written to means losing data. So nothing runs until a three-part write-stop gate agrees:

  1. An external signal that the source is ready to stop.
  2. A machine check that the database really is write-free.
  3. An operator’s explicit approval.

Any one of those alone can be wrong. All three together are hard to fool.

Steps you can rerun

The heavy lifting runs as PowerShell 7 runbooks on dedicated “middleman” VMs, dispatched through Azure Automation. Every step is written to be idempotent: running it twice has the same effect as running it once. That property is what makes “resume from the failed step” safe instead of scary.

Prep scripts work the same way: a manifest of small, selectable scripts with dependency rules, validation for each, and resume at the exact script that failed.

Zero stored credentials

Migrations touch personal data, so the security model can’t be an afterthought:

  • Managed identities and Key Vault throughout, with no passwords in config files or pipelines.
  • Private endpoints and an EU region, because GDPR applies.
  • Operator-only access through Microsoft Entra ID with a dedicated role.

Test it without the cloud

My favourite part: the whole pipeline runs in a simulation mode on mocks, with no Azure at all. Forced failures, mitigation, resume, cancellation, cutover and cleanup can all be exercised on a laptop. The riskiest logic gets tested the most, and for free.

Takeaways

  • Model long operations as persisted state, not as a script.
  • Make every step idempotent so resuming is safe.
  • Gate the dangerous start on independent checks.
  • Keep the final switch human.
  • Simulate the failure paths, not just the happy one.

Get in touch

Tell me about your project, role or idea. I usually reply within a couple of days.

I only use your name, email and message to reply to you. No cookies, no tracking.Privacy policy

Esc