This workflow covers a Linux server running an agent application that calls a model API. It is not a GPU sizing guide or a production-ready installation script. It makes decisions and checks visible before you create a paid resource.
We reviewed the linked documentation but have not executed this workflow end to end. No instance was purchased or deployed for this article. Interface labels and available images may change. Use the current official instructions for your selected product and operating system; do not translate this checklist into commands without understanding the effect of each one.
1. Write the deployment record first
Record the application release, expected concurrent tasks, persistent data, administrator, and monthly spending envelope. Name the model provider and the external tools the application can use. Decide whether this is a disposable experiment or a service whose data must survive failure.
Choose a region with the application’s users and dependencies in mind. A nearby region is a candidate for testing, not proof of acceptable latency or compliance. Choose a supported Linux release you can maintain, then check that the application’s runtime supports it. Avoid picking an image only because a tutorial uses the same screenshot.
Select compute, memory, and disk based on the workload you will measure. Leave headroom for updates, logs, and temporary document processing. Review the current configuration and recurring charges, including optional services, before deploying. Vultr’s Cloud Compute provisioning guide documents the available decision points, including image, SSH key, firewall group, and networking options.
2. Prepare access before you need it
Create or select an SSH key on a trusted administrator device using that operating system’s current instructions. Keep the private key on the administrator device; the server and provider need the public key. Give the key a descriptive label so you can identify and remove old access later.
Select the intended public key during provisioning. Record the instance identity and the recovery-console route. On first connection, verify the server’s host identity through a trusted route rather than blindly accepting an unexpected host-key warning. Test the intended non-root administrator account and its required privileges before removing any existing access method.
Do not confuse changing a key record in the provider account with updating an existing server. Vultr’s SSH key guide warns that the console’s reinstall-based key workflow wipes the server. For an existing machine, follow the documented in-system key-management path and verify a second working session before closing the first.
3. Design the network before exposing the app
Make an explicit port inventory. For a typical web application, the public entry point is an authenticated HTTPS service. Administrative access should be restricted to your known access path. The database, queue, model gateway, and internal application listener generally do not need public exposure.
A Vultr firewall group applies rules to attached instances and supports IPv4 and IPv6. Confirm that the correct group is actually attached; creating a group alone does not protect an unrelated instance. Use the firewall-group reference for the current controls.
Review the host firewall as a separate layer. Preserve an allowed administration path before enabling restrictive rules, and make sure you can recover from a mistake. Vultr’s firewall troubleshooting guide distinguishes provider and instance-level filtering. Do not copy a diagnostic “disable firewall” step into a permanent configuration.
4. Keep internal container ports internal
Decide whether the reverse proxy runs on the host or inside the container network. Bind services to the appropriate private interface or internal network, and publish only the ports that the design requires. Avoid exposing a database port for convenience when a private administration route would suffice.
Docker’s firewall documentation explains that published container traffic can be diverted before ufw’s filtering rules. An “active” ufw status is therefore insufficient proof that a published port is unreachable. Understand your Docker network mode and verify the actual network boundary from outside the machine.
Install the chosen application from a trusted, documented source. Review its configuration, secret handling, volumes, and tool permissions before adding real credentials. Keep model keys server-side. Start with non-sensitive test data and tools that cannot perform consequential actions.
5. Collect launch evidence
Before calling the server ready, record the outcome of these checks:
- A fresh administrator session succeeds through the intended route
- The public hostname serves the correct application with valid HTTPS
- An unauthenticated visitor cannot reach administrative functions
- Internal application, database, and queue ports are unreachable externally over every enabled address family
- A harmless test task completes, while an intentionally forbidden tool action is denied
- Restarting the application preserves the state you intended to keep
- Logs omit credentials, monitoring reaches its owner, and a recovery exercise succeeds
Do not mark a check as passed because a configuration file appears correct. Test the behavior and keep enough evidence to repeat the check after a change. If you cannot test a property yet, record it as unverified and withhold the access that depends on it.
6. Plan the end of the experiment
Set a review date and list everything the experiment creates. If you abandon it, preserve only the data you need, verify the export, and remove resources deliberately. On Vultr, stopping an instance does not stop its billing; destruction deletes the instance’s data. Review the billing policy before cleanup.
Check separately for retained storage, snapshots, and other subscriptions. A clean shutdown is an operational event. A clean teardown is also a data-management and billing event, and deserves its own verification.