Moving an agent is more than copying files and updating a hostname. Background schedules, queued jobs, callback endpoints, and saved credentials can keep working after you think traffic has moved. A migration succeeds when the new system owns the work, the old system cannot create conflicting actions, and you know which state is authoritative.
This is an unexecuted reference workflow. We have not migrated a live agent or measured downtime for this article. The steps need adaptation to your database, queue, application, and availability requirements. A single-server service can often accept a short maintenance window; do not promise zero downtime without a design and tests that support it.
Define the move and the exit criteria
Write down why you are moving: capacity, cost, location, operational control, or a provider change. List what is staying the same. Changing the model, database engine, application version, and hosting provider simultaneously makes failures harder to isolate.
Inventory every resource: images, durable files, database, DNS, scheduled jobs, inbound webhooks, outbound integrations, secrets, monitoring, and backups. Identify hard-coded addresses and provider-specific dependencies. Record which person can authorize cutover and what evidence they need.
Set explicit success criteria, such as correct authentication, intact representative records, successful harmless tasks, and normal queue processing. Define rollback triggers and a latest safe decision point before beginning. Make sure users know about a planned interruption when the service requires one.
Rehearse the transfer
Build the destination with the intended release and configuration, but keep scheduled workers and consequential tools disabled. Transfer a recovery copy and test it in isolation. Docker’s volume documentation distinguishes persistent data from container lifecycle; moving an image alone does not migrate that data.
Use the database’s supported export, replication, or backup method. PostgreSQL’s backup overview, for example, distinguishes approaches with different recovery assumptions. Check compatibility instead of assuming that a raw data directory can be opened by any newer database image.
Time the rehearsal. Include data transfer, verification, certificate readiness, and operator actions. If it takes longer than the allowed maintenance window, revise the plan before the real cutover. The rehearsal is also where you discover missing plugins, file ownership issues, or credentials that were never documented.
Choose exactly who can write
For a simple maintenance-window migration, stop admitting new work, pause schedules, and drain or explicitly checkpoint active tasks on the source. Record which tasks completed, which are still pending, and which may have performed an external action without recording success.
Transfer the final consistent state only after the write boundary is established. Keep the old worker disabled while enabling the new one. Two copies of an agent using the same mailbox, queue, or schedule can both act, even if only one receives website traffic.
For every uncertain action, check the downstream system before replaying it. Preserve task identifiers and idempotency records where the integration supports them. A queue that looks empty is not sufficient evidence that no worker still has a task in memory.
Cut over traffic and observe
Verify the destination through an appropriate test route before changing public routing. Confirm its identity, TLS behavior, authentication, and internal-service boundaries. Update only the records or routing controls in the approved plan.
DNS caches can continue directing clients to the previous address. Cloudflare’s TTL documentation explains the caching trade-off. Plan for overlap rather than treating a DNS edit as an atomic switch. The old endpoint should have a deliberate safe behavior, such as maintenance mode or a controlled forwarding arrangement, without restarting its worker.
Watch application errors, authentication failures, task throughput, callback delivery, and duplicate-action reports. Test from outside your own network. A working administrator session may hide problems affecting ordinary users or clients with cached DNS.
Make rollback state-aware
Before the destination accepts writes, returning traffic to the source may be straightforward. After new writes or external actions occur, the source is stale. Do not simply point DNS back and resume both systems.
Pause new work, reconcile the authoritative state, and follow the rollback plan appropriate to those changes. If safe rollback cannot preserve newly accepted work, communicate that limitation and choose a controlled recovery path. Keep the source available for the agreed observation period, with its action-producing components disabled.
Retire resources deliberately
Once the exit criteria are met, take the agreed final archive, verify its retention owner, and revoke old system access. Remove obsolete callback URLs, monitoring targets, scheduled jobs, and DNS records. Review remaining storage and backups separately from compute.
Vultr’s server billing policy states that powered-off servers remain billable until destroyed. Keeping a rollback server therefore has a cost. Record its retirement date and verify resource removal after the retention decision is approved.
Finish with two checks: the new service owns all intended work, and the old project has only the resources and charges you deliberately retained. Migration is complete when both operational control and the resource ledger agree.