Server Manager
The Server Manager plugin provides a web UI for managing remote Joinery production servers. Operations include status checks, backups, database copies, and applying updates -- all from the admin interface at /admin/server_manager.
The system has two components:
- PHP plugin (
plugins/server_manager/) -- admin UI, job creation, command generation - Go agent (
/home/user1/joinery-agent/) -- generic step executor that polls the job queue and runs commands via SSH
Quick Start
1. Install the plugin
The plugin is already in the plugins/ directory. From the admin panel:
- Go to
/admin/admin_plugins - Click Actions on "Server Manager" and choose Install
- Click Actions again and choose Activate
mgn_managed_nodes, mjb_management_jobs, ahb_agent_heartbeats, bkt_backup_targets.2. Install and start the Go agent
Release channel (how the agent normally arrives and stays current)
The agent ships inside the platform release. Publishing an upgrade bundles a signed agent artifact into plugins/server_manager/agent_dist/:
manifest.json— agent version plus, per architecture, the artifact filename, its sha256, and an Ed25519 signature over the raw binaryjoinery-agent-linux-amd64.gz/joinery-agent-linux-arm64.gz— the binariesjoinery-agent.service— the systemd unit
publish_upgrade.php cross-compiles both architectures from the checkout named by the server_manager_agent_source_path setting (default /home/user1/joinery-agent) whenever the source version differs from the bundled one, and signs them with the key at {site root}/config/agent_signing_key (generated on first publish; the .pub sibling holds the base64 public key that gets baked into the built agent).Bundling is the first thing a publish does, before the VERSION file, the archives or the release row, because its outcome decides whether the release happens at all:
- No agent source on this box — the existing artifact carries forward unchanged and the publish proceeds. Publishing never depends on a Go toolchain being present.
- Source version matches the bundle —
agent_distis left byte-identical, which keeps theserver_managerplugin tree hash stable between publishes. - Source is newer and the rebuild succeeds — the fresh artifact is bundled, and because this happens before plugin archives are built it is captured in the
server_managerarchive and its tree hash. - Source is newer and the rebuild fails — the publish is refused. The build error is printed, and the VERSION file, archives and release row are all left untouched. Shipping here would mean releasing an agent the publisher already knows is out of date, and the resulting fleet has no way to tell.
plugins/server_manager/tests/agent_bundle_drift_test.php asserts the same invariant on its own, so a bundle that falls behind its source is caught by the safe test tier rather than by the next release.First install is handled by the plugin's host_installer (provisioning/install_agent.sh), which runs at every root moment — site install, code upgrade, container start, and the node-detail Run Plugin Installers action. It installs the bundled binary, writes the env file with the right JOINERY_CONFIG, and sets up systemd or cron supervision automatically.
Every later version change is handled by the agent itself. Between jobs, the agent compares its own version with the bundled manifest. When they differ, it decompresses the artifact, checks the sha256, verifies the Ed25519 signature against the public key embedded in its binary, keeps the current binary as .bak, renames the new one into place, and exits cleanly for its supervisor to restart. The signature check is the security boundary: the site tree is writable by the web user while the agent runs as root, so the agent never installs anything the publisher did not sign. An artifact that fails verification is refused, logged under a === Self-update === header, and not retried until the manifest changes.
If the new binary fails to initialise (config, DB, or schema), it restores the .bak over itself and records the bad version in a .rejected marker — the supervisor restarts the previous working agent, and that version is never reinstalled; the next release supersedes the rejection. On the first fully healthy start after an update, the .bak and any stale marker are removed.
The dashboard's Agent Status bar surfaces all of this from the heartbeat row (ahb_bundled_version, ahb_update_state): a pending update, a refused (verification-failed) artifact, a rolled-back version, or an agent built without an update key.
Manual install (bootstrap fallback)
For a control plane that has no bundled artifact yet, build and install by hand:
cd /home/user1/joinery-agent
PUBKEY=$(cat /var/www/html/joinerytest/config/agent_signing_key.pub) make release
sudo bash joinery-agent-installer.sh --verbose [--config /path/to/Globalvars_site.php]make release compiles the binary (passing PUBKEY bakes in the update-verification key so the manual build can still self-update later) and packages it into joinery-agent-installer.sh, a self-extracting script that handles both fresh installs and upgrades with automatic rollback if the new version fails to start.
The installer detects the host's supervision capability:
- systemd hosts: installs
/etc/systemd/system/joinery-agent.service; start withsystemctl start joinery-agent. - No systemd (Docker containers, minimal hosts): installs
/usr/local/bin/joinery-agent-superviseplus/etc/cron.d/joinery-agent(@reboot+ a once-a-minute keepalive) and starts the agent immediately. Logs go to/var/log/joinery-agent.log.
/usr/local/bin/joinery-agent— the binary/etc/joinery-agent/joinery-agent.env— configuration (from example, first install only)
--config PATH stamps JOINERY_CONFIG into the env file, pointing the agent at the right site without editing anything.Every control plane needs a live agent. All jobs (install_node, provision_ssl, backups, upgrades) sit pending until an agent polling that site's own database claims them. The Server Manager → Provisioning page shows an agent heartbeat badge as requirement #1.
Configure (usually not needed)
The agent reads database credentials directly from Globalvars_site.php — no manual configuration required on a standard Joinery install.
The default config path is /var/www/html/joinerytest/config/Globalvars_site.php. If your install is at a different path, set it in the env file:
sudo nano /etc/joinery-agent/joinery-agent.env
# Set: JOINERY_CONFIG=/var/www/html/mysite/config/Globalvars_site.phpOther optional settings in the env file:
| Setting | Default | Purpose |
|---|---|---|
JOINERY_CONFIG | /var/www/html/joinerytest/config/Globalvars_site.php | Path to Globalvars_site.php |
POLL_INTERVAL | 5s | How often to check for new jobs |
HEARTBEAT_INTERVAL | 30s | How often to update the dashboard status |
AGENT_NAME | joinery-agent | Name shown in the admin dashboard |
DB_HOST, DB_PORT, DB_NAME, DB_USER, DB_PASSWORD | (from Globalvars) | Override DB credentials if needed |
Start
sudo systemctl start joinery-agent
sudo systemctl status joinery-agentThe dashboard at /admin/server_manager should now show Agent Status: Online.
If anything is wrong, the agent logs to the systemd journal (systemd hosts) or /var/log/joinery-agent.log (cron-supervised hosts):
journalctl -u joinery-agent -f # systemd
tail -f /var/log/joinery-agent.log # cron supervisionCommon startup errors are self-explanatory — missing DB_NAME, wrong password, or plugin tables not installed. Each error message tells you exactly what to fix.
Upgrade
Agents upgrade themselves from the bundled artifact after each platform release lands (see Release channel above). The manual-install path also accepts upgrades: re-running the generated installer stops the service, swaps the binary, restarts, and rolls back automatically if the new version fails to start.
3. Add managed nodes
Go to /admin/server_manager/node_add (or click Add Node on the dashboard). There are two ways to add nodes:
Auto-detect (recommended)
The Auto-Detect Joinery Servers panel scans a remote host for Joinery instances automatically. Enter:
- SSH Host -- the server IP (e.g.,
23.239.11.53) - SSH Key Path -- path to the private key on the control plane (defaults to
/home/user1/.ssh/id_ed25519_claude) - Click Detect
discover_nodes job. The Go agent SSHes to the host, finds Docker containers (or bare-metal installs) running Joinery, and reports back with each instance's container name, web root, domain, database name, and version.Detected instances appear as cards with Add This Node buttons. Clicking one auto-fills the entire form below -- just click Add Node to save.
Auto-detect requires the Go agent to be running (it executes the SSH commands, not PHP).
Manual
Fill in the form fields directly:
| Field | Example (Empowered Health) | Example (ScrollDaddy) |
|---|---|---|
| Display Name | Empowered Health Production | ScrollDaddy Production |
| Slug | empoweredhealthtn | scrolldaddy |
| SSH Host | 23.239.11.53 | 23.239.11.53 |
| SSH User | root | root |
| SSH Key Path | /home/user1/.ssh/id_ed25519_claude | /home/user1/.ssh/id_ed25519claude |
| SSH Port | 22 | 22 |
| Docker Container | empoweredhealthtn | scrolldaddy |
| Container User | (blank) | (blank)_ |
| Web Root | /var/www/html/empoweredhealthtn/public_html | /var/www/html/scrolldaddy/public_html |
| Site URL | https://empoweredhealthtn.com | https://scrolldaddy.app |
Admin Pages
All pages are at /admin/server_manager/... and require permission level 10 (superadmin).
The UI is organized around a dashboard + node detail pattern. The dashboard shows the fleet overview; clicking a node opens a tabbed detail page with all operations for that node.
| URL | Purpose |
|---|---|
/admin/server_manager | Dashboard -- agent status, node cards with health dots, publish upgrade, recent jobs |
/admin/server_manager/node_detail?mgn_id=N | Node Detail -- tabbed page for a single node (see tabs below) |
/admin/server_manager/node_add | Add Node -- auto-detect panel + manual add form |
/admin/server_manager/targets | Backup Targets -- CRUD for cloud storage targets (B2, S3, Linode) |
/admin/server_manager/jobs | Jobs -- global job history with filters by node, status, and type |
/admin/server_manager/job_detail?job_id=N | Job Detail -- single job output with live polling |
Node Detail Tabs
The node detail page (/admin/server_manager/node_detail?mgn_id=N&tab=...) has six tabs:
| Tab | Purpose |
|---|---|
| Overview | Status summary (health dot, disk/memory/load/postgres/version), action buttons (Check Status, Test Connection), recent jobs for this node, connection settings (collapsed by default), delete node. The Actions dropdown also offers Run Plugin Installers — queues a run_plugin_installers job that executes every active plugin's declared host_installer on the node as root (idempotent); this is how a bare-metal node picks up system-service configuration (e.g. the mail stack) after a plugin is activated, since it has no container-start moment |
| Backups | Target indicator, run database/project backup, backup file browser with scan, per-file upload-to-cloud and delete, restore full project from a .tar.gz archive, restore from an incremental chain |
| Database | Copy database from another node to this one, restore from backup file |
| Updates | Version comparison (node vs control plane), apply update |
| Jobs | Job history filtered to this node, with status and type filters |
| Console | Run one ad-hoc command on this node. Off until the node opts in — see below |
Node console
Some fleet work needs a command no built-in job covers. The Console tab runs one, on one node, and records it.
Turning it on. Overview → Edit → Allow running commands from the Console
tab (mgn_allow_console, off by default). Until then the tab shows only a link
back to that setting. A newly added or provisioned node is never
console-reachable by default.
Running a command. Type it, choose how long to wait (1, 2, 5 or 10 minutes), and — on a container node — choose whether it runs inside the container or on its host. A confirmation dialog shows the resolved execution context (which SSH user, which host or container, the timeout, the command) before anything is sent, so what is confirmed is what actually runs.
The gate. A command is refused unless all of these hold: the operator is a
superadmin (permission 10), the node has mgn_allow_console on, and the
operator has confirmed a second factor recently in this session — the standing
step-up rule in Account Security. When a
confirmation is owed, the passkey ceremony runs inline and the form then
submits. An account holding no second factor passes that check; the gate binds a
factor the account has rather than requiring an enrollment. A refusal
re-renders the tab with the typed command intact.
What is not checked is the command. Nothing inspects, classifies or filters it — no shell string can be judged safe by inspection, and a filter that implied otherwise would be worse than none. The bounds are the gate above, the timeout (which kills the command and fails the job), and the audit record.
Privilege is the node's SSH identity's: mgn_ssh_user over the node's key.
If that user has no sudo, neither does the console.
The record. Every run is a run_command management job — command, operator,
node, timestamps, exit status and captured output — on the jobs pages, on the
job detail page with its live output stream, and in the tab's own Recent
commands list. node_exec.php writes the same row (marked as coming from the
CLI, with no operator), so both ways onto a node produce one audit stream.
Stored commands and output are redacted for credential material, including
PGPASSWORD=-style env prefixes.
Dashboard Features
The dashboard shows:
- Agent Status -- online/offline indicator with version and last heartbeat time
- Managed Nodes -- cards with health-based status dots (green=healthy, yellow=warning, red=problem, gray=no data), key metrics, and action buttons
- Publish Upgrade -- build upgrade archives from control plane source code (node-independent)
- Recent Jobs -- latest 20 jobs across all nodes
- Red: Last check failed, disk > 90%, or PostgreSQL not accepting connections
- Yellow: Disk > 80% or load average > 5
- Green: All metrics healthy
- Gray: Never checked or no data
Job Types
| Job Type | Description | Destructive |
|---|---|---|
check_status | SSH-probe disk, memory, uptime, PostgreSQL, version; subsumes the old test_connection since its first step is the SSH handshake | No |
backup_database | Run backup_database.sh, optionally upload to cloud | No |
backup_project | Run backup_project.sh (DB + files + Apache config), optionally upload | No |
list_backups | List backup files on local server and cloud target | No |
upload_backup | Push one existing backup file from the node to its cloud target; keeps the local copy | No |
delete_backup | Delete backup files from local, cloud, or both | Yes |
copy_database | Dump source DB, transfer, restore on target | Yes |
restore_database | Restore a backup file on a node | Yes |
restore_project | Restore a full project .tar.gz (files + DB) in place on an existing node, then reconcile it to that machine. Runs restore_project.sh --force --domain <domain>, which cascades --non-interactive into restore_database.sh. Pre-restore snapshots of DB and files written to /backups/auto_pre_project_restore_*. Every file in the archive must exist under the project directory afterwards or the restore fails and names what is missing | Yes |
restore_chain | Restore a node from an incremental backup chain — what the fleet's scheduled backups actually produce. Fetches the chain manifest, recovers the chain key on the node from the node's own backup_site_key, downloads every artifact the manifest names up to the chosen run, then runs restore_chain.sh, which verifies each artifact against its recorded size and hash before writing anything and applies them in order | Yes |
apply_update | Run upgrade.php on target | Yes |
publish_upgrade | Run publish_upgrade.php locally on control plane (in plugin) | No |
discover_nodes | Scan a remote host for Joinery instances (Docker + bare metal) | No |
install_node | Provision a fresh Joinery site on a remote host (fresh or from-backup) | No (target must be clean) |
provision_ssl | Run certbot on the node's host to obtain a Let's Encrypt cert | No |
backup_run | This control plane's own backup of a node. The node runs its backup engine — chain, envelope, upload, local sweep — with the bucket, a write-only credential and the recovery key supplied for that run and never stored there | No |
run_command | One ad-hoc command from the node detail Console tab (or node_exec.php). Bounded by a chosen timeout; the command itself is never inspected | Depends entirely on the command |
decommission_node | Ship and run remove_account.sh on the host to permanently delete the site, verify it is gone, then soft-delete the node record | Yes |
Note on bare-metal nodes with user1 SSH: When a bare-metal install completes, install.sh disables root SSH and the node's mgn_ssh_user is automatically updated to user1. Subsequent jobs run as user1 with NOPASSWD sudo. All backup/restore commands that need root-level paths (e.g. /backups/) use sudo automatically.
One-Click Node Install
Dashboard → Install New Node opens a form that provisions a Joinery site in a single click. The Target Host dropdown offers three kinds of target:
- A known host (or Other server with manual SSH details): the form creates the ManagedNode and dispatches the
install_nodejob immediately. - Create a new cloud instance: no server exists yet. The form records an admin-origin
CustomerCloudProvision(connected cloud account, region, instance type, plus all install parameters) and the Provision Customer Cloud task births the instance, creates the node, and dispatches the install — see Customer-Cloud Fulfillment. The instance is created in, and billed to, the selected connected account; Linode grants expire after two hours, so connect (or re-connect) shortly before submitting. Cloud targets always take a fresh source backup in From-Backup mode. In-flight provisions appear in a banner at the top of the dashboard.
mgn_skip_joinery_checks set — but no site is installed (no web root, site URL, or SSL flow). Completion is a passing check_status job. This is how infrastructure nodes that host no Joinery site — a mail relay shard, for example — enter management; the role's own provisioning (e.g. the mailbox plugin's provision-relay job) builds on the bare node afterward. Bare is admin-origin only; orders always install a site.Two install types:
- Fresh: empty Joinery site with default schema. Admin picks the domain. The admin login is
[email protected]with a password generated for that site alone — there is no shared default. Like the generated Postgres password, it stays on the node rather than in the control plane: read it at/var/www/html/{sitename}/config/admin_credentials.txt(root only), or set a new one withmaintenance_scripts/sysadmin_tools/reset_admin_password.php.usr_force_password_change=true, so the first sign-in forces a new password. - From Backup: fresh install + restore of a source node's DB and project files, then reconciliation to the new node — its own domain (the node's recorded URL), its own deployment shape, its own paths. Use source admin credentials to log in; cut DNS over when ready and the certificate is issued on its own.
maintenance_scripts/install_tools/ are packaged locally, SCP'd, extracted on the target, and install.sh -y -q site SITENAME - DOMAIN runs non-interactively. Docker installs add a follow-up step that invokes manage_domain.sh set SITENAME DOMAIN --no-ssl on the target to auto-install Apache + mod_proxy (if missing) and wire up an HTTP reverse proxy on port 80 — so the site is reachable at http://DOMAIN/ as soon as DNS points here. SSL stays a separate admin step (certbot --apache -d DOMAIN on the target). For From-Backup, source backups are captured (or an existing cached backup is used), fetched to the control plane, and pushed to the target after install.From-Backup restores files by extracting the source archive with both of its
leading path components stripped, taking only the project_files/ subtree —
backup_project.sh writes archives as `{backup_name}/project_files/{public_html,
uploads,config,...} with the archive's own metadata (apache_config/`,
backup_info.txt, the .sql dump) as siblings. The target keeps its own
Globalvars_site.php (it holds this machine's database password and
secret_box_key) and mints its own backup_site_key rather than inheriting the
source's identity as a backup recipient. A verification step then requires every regular file the
archive carries to exist at the site root and fails the job otherwise, because a
clone whose files did not land still serves pages: the fresh install ran first
and the database restore succeeded, so the only symptom is uploaded files
missing from where the restored database says they are.
The mgn_install_state column tracks the lifecycle: installing → NULL (success) or install_failed (failure). On failure, the node detail page surfaces a Retry Install button; the target must be cleaned manually (e.g. rm -rf /var/www/html/SITENAME) before retry because install.sh refuses to overwrite an existing site. Postgres passwords are auto-generated and stored in the target's Globalvars_site.php — Server Manager does not capture or display them.
Docker notes:
- The reverse proxy step (
manage_domain.sh) is skipped when the domain is a bare IP address — a routable hostname is required for ApacheServerName-based virtual hosting. With an IP domain, the site is accessible directly on its mapped port. backup_project.shrequiresrsync. The bare-metal and Docker install scripts install rsync as part of the essential packages (install.shline ~948). Sites installed before this was added can install it manually withapt install rsync.- After a Docker install,
mgn_container_nameis automatically recorded in the control plane DB so future jobs correctly usedocker execto reach the site.
SSL Management
SSL State
Each node tracks its TLS certificate state in mgn_ssl_state:
| Value | Meaning |
|---|---|
null | Unknown or not configured |
pending | Waiting for DNS propagation; certbot has not run yet |
active | A valid Let's Encrypt cert is installed |
failed | Provisioning failed after repeated retries |
Automatic Detection
check_status jobs include an SSH step that checks for a Let's Encrypt cert under /etc/letsencrypt/live/{domain}/. JobResultProcessor updates mgn_ssl_state and stores ssl_domain, ssl_expiry_raw, and ssl_expiry_ts in mgn_last_status_data. State transitions:
CERT_FOUND→ sets state toactive(from any prior state)CERT_MISSING→ clears state tonullonly if currentlynulloractive; never overwritespendingorfailed
Manual Provisioning
The Overview tab shows an SSL Setup card when mgn_ssl_state is not active, the node has a domain in its site URL, and mgn_cert_expiry_ts is empty. The last condition excludes directly-exposed, self-renewed nodes (see Certificate Expiry Monitoring) — their cert lifecycle is owned by an external renewer (e.g. Caddy), and the card's certbot-based provisioning does not apply to them. The card:
- Resolves the domain via DNS and shows whether it points to the node's host IP
- Enables the Provision SSL button when DNS is ready (or when the host IP is not configured)
- On submit: creates a
provision_ssljob, setsmgn_ssl_state = 'pending', redirects to job detail
provision_ssl job runs certbot --apache -d DOMAIN on the node's host (for Docker nodes, certbot runs on the reverse-proxy host, not inside the container). On success, mgn_ssl_state is set to active by JobResultProcessor.Cloudflare-proxied domains skip certbot (Cloudflare terminates TLS at its edge) but are gated on a routing probe: the job writes a one-time token into the site's webroot, and the control plane fetches it through the domain. Only a match — proof that traffic for the domain actually lands on this node — patches the proxy's X-Forwarded-Proto and marks SSL active (JobResultProcessor additionally requires the CF_ROUTING_VERIFIED marker). A miss fails the job and the domain stays pending until the customer's DNS actually routes here.
Automated Provisioning (installs only)
For nodes installed via Install New Node, ProvisionPendingSsl (scheduled hourly) watches for nodes with mgn_ssl_state = 'pending', checks DNS, and kicks off provision_ssl jobs automatically. After ~16 hours of failed attempts it flips state to failed — except a Cloudflare domain still waiting on its DNS cutover (CF_ROUTING_UNVERIFIED), which keeps quietly retrying. Manual provisioning via the Setup card is the fallback.
Hosting Provisioning
Paid hosting orders on getjoinery become installed, SSL'd Joinery sites with
no human touch. The Poll Hosting Orders scheduled task polls the
getjoinery API each cron tick for paid orders carrying an answer to the
configured domain Question, and fulfills each one in the mode the product
declares in pro_fulfillment_provider:
- Shared host (the default, any other value): the pipeline picks the
least-loaded provisioning-enabled
ManagedHost, assigns the next Docker port, and dispatches aninstall_nodejob. The buyer's site is a container on infrastructure the operator owns. customer_cloud: the buyer's site runs on a server in *their own cloud account, billed to them by the provider — see below. A product opts in by picking Customer cloud server in the product-edit Purchase grants picker (CustomerCloudFulfillment, registered with the store's FulfillmentRegistry from serve.php); that stamps the provider value and contributes the domain question as a checkout requirement automatically.
install_node completes, the welcome email goes
out with DNS instructions, and ProvisionPendingSsl turns HTTPS on once DNS
resolves.Activation — the Provisioning page
Server Manager → Provisioning (/admin/server_manager/provisioning_setup)
activates the pipeline: every requirement shows a live status badge, and each
automatable step is a one-click, idempotent action backed by
includes/ProvisioningSetup.php — mint the store API service user
(provisioning@<host>, permission 5, password recovery disabled) and machine
key and write the API settings (with a loopback probe badge and key
rotation), create the domain Question, save the email settings, activate the
three scheduled tasks, and save the customer-cloud settings (SSH key path
with key/.pub existence badges, referral URL, instance defaults). The page
also shows what stays manual: attaching the question to hosting products,
opting a shared host in, and registering the Linode OAuth app. When the store
is a remote site rather than the control plane itself, the service key is
minted on the store site and its values entered in the API settings fields.
The customer-cloud provisioning keypair (the public half is installed on
created instances; the private half is the control plane's only access to
them) is generated automatically at plugin activation
(activate.php → ProvisioningSetup::ensureSshKey()), defaulting to
{site root}/config/provisioning_key. The page's Generate provisioning
key button runs the same idempotent action for control planes activated
before the key existed; an existing key or custom path is never overwritten.
Customer-Cloud Fulfillment
The buyer connects their Linode account once at
/profile/server_manager/connect_cloud (the Connect page — also the
re-connect page if a grant is later revoked). The grant flows through the
platform OAuth2 core (provider linode, consumer purpose
customer_cloud, scope linodes:read_write only — no account or billing
access). Tokens are SecretBox-encrypted on the buyer's
CustomerCloudAccount row.
Each provision is a CustomerCloudProvision row that the
Provision Customer Cloud scheduled task advances:
pending_connect → (grant arrives) → ready → instance created on the
connected account → booting → running + IP → ManagedNode + install_node
job → installing → done (or failed, which alerts the ops address).
Provisions have two origins (cvp_origin):
- order — created by a customer-cloud purchase. Starts at
pending_connect, installs fresh + Docker, and sends the buyer welcome email on completion (the order-item linkage drives it). - admin — created by the Install New Node form's cloud-instance target.
Starts at
ready(the admin picked an already-connected account), carries its install parameters on the row (cvp_docker_mode,cvp_install_mode,cvp_source_node_id,cvp_backup_source,cvp_port,cvp_sitename), and sends no welcome email. The row belongs to the grant owner (cvp_usr_user_id), so a stale grant is re-connectable by the person who can actually re-consent.
server_manager_linode_referral_url
so new Linode signups carry the referral credit. A token-refresh failure or a
provider 401 parks the provision back at pending_connect and flags the
account link — a fresh grant resumes it automatically.Customer-owned node semantics: the resulting node is a normal
ManagedNode (installs, upgrades, uptime checks, SSL all apply), with
mgn_ssh_key_path set from server_manager_customer_cloud_ssh_key_path
(whose .pub sibling is installed on the instance at create time) and no
mgn_mgh_host_id — it belongs to no managed host. The server is the
customer's property: cancelling their subscription stops management, never
touches the instance.
Settings: server_manager_customer_cloud_ssh_key_path (required),
server_manager_customer_cloud_region / _type / _image (instance
defaults), server_manager_linode_referral_url. Provider credentials are the
core oauth_linode_* settings (Admin → System → OAuth Providers).
The compute API surface is CloudComputeProvider
(includes/cloud_compute/) with LinodeComputeDriver implementing it; a new
provider is a new driver plus its OAuth provider class.
Backup Targets
Backup targets define where backup files are uploaded after creation. Each node can optionally have a backup target assigned. If no target is set, backups remain local only on the remote server.
Supported Providers
| Provider | Credentials (UI fields) |
|---|---|
| Backblaze B2 | Application Key ID + Application Key (region/endpoint auto-detected via b2_authorize_account at save time) |
| Amazon S3 | Access Key + Secret Key + Region |
| Linode Object Storage | Access Key + Secret Key + Region + Endpoint URL |
S3Signer.php. There is no per-provider CLI dependency — uploads, downloads, deletes, and listings all run as direct HTTPS calls from either the control plane (web tier) or the node (via a heredoc'd node_uploader.php script). New S3-compatible providers can be added by configuration alone, no script changes.Nodes with no backup target leave backups local-only on the remote server.
Configuration
- Go to
/admin/server_manager/targetsand click Add Target - Select a provider, enter bucket name, path prefix, and credentials
- Go to a node's Overview tab, expand Edit Connection Settings, and select the target from the Backup Target dropdown
- Save — backups for this node will now auto-upload after creation
Upload Path Structure
All providers use: {prefix}/{node_slug}/{filename}
Example: joinery-backups/empoweredhealthtn/empoweredhealthtn-04_11_2026.sql.gz.enc
Credential Storage
Credentials are stored on the bkt_backup_targets table using a unified shape for every provider:
{"access_key": "...", "secret_key": "...", "region": "...", "endpoint": "..."}Two columns hold two keys: bkt_credentials is the main (delete-capable) credential the control plane itself uses, and bkt_node_credentials optionally holds a write-only key handed to nodes instead (see The node may write to the shelf but never erase it). Both are SecretBox-sealed at rest.
A persisted job command never contains a credential — it carries a placeholder token that the agent resolves in memory immediately before the step runs: __SM_CREDS_<target_id>__ for the main slot, __SM_NODE_CREDS_<target_id>__ for the node slot. The builder decides at build time which token a step gets: node-side uploads carry the node token whenever the node slot is filled, while node-side downloads and cloud deletes always carry the main token, because those need the read and delete capability a write-only key deliberately lacks. The agent resolves exactly the slot the token names and never falls back to the other, so a job built against a since-emptied slot fails visibly rather than running with a more powerful key than intended.
For node-side operations (upload, delete, download), the resolved credentials are embedded into a self-contained PHP script that is piped to the node via a heredoc'd php -- invocation — never written to a file on the node and never visible in process listings as positional arguments. The S3Signer.php and node_uploader.php source is composed at job-build time by JobCommandBuilder::build_node_uploader_script().
Because the script is composed from the control plane's own copy of those two files, changes to the signer or the uploader reach every node on its next job — there is no agent release or node upgrade in the loop.
Transient Failures
A storage provider that answers a request with a 5xx does not fail the job. S3Signer::request() retries — MAX_ATTEMPTS tries, exponential backoff with jitter — for the failures that are worth another go: 5xx, 429, 408, and the transport errors that mean the connection died rather than the request being wrong. A deterministic error (403 signature, 404, 400) is returned immediately; retrying it would only burn the budget and bury the message. S3Signer::is_retryable() is the whole policy and is pure, so the classification is testable without a network.
Retrying is safe because every request the class makes is idempotent: a PUT overwrites its key (a single PUT, never multipart, so no orphaned parts survive a failure), and GET/DELETE/list have no cumulative effect.
Two bounds keep a retry from doing harm:
- Wall clock. Total time is capped at one attempt's timeout plus
RETRY_WINDOW_SECONDS. An attempt that burns the entire transfer timeout leaves no room for another — the right answer, since a transfer that cannot finish in an hour will not finish on the second try. Job steps that shell out to the uploader take their owntimeoutfromS3Signer::transfer_budget_seconds(), so the agent can never kill a transfer part-way through a retry. - Replay. A retried upload rewinds the body stream before resending. curl consumes the stream on the first attempt while
CURLOPT_INFILESIZEstill claims the full length, so a retry without the rewind sends nothing and then blocks until the timeout — a hang rather than an error. A stream that cannot seek is not retried at all, rather than being sent truncated.
RETRY: attempt 1 failed (HTTP 500 internal incident); retrying in 2s), and a transfer that only succeeded on a later attempt says so. A provider that is degrading looks exactly like a healthy one unless the attempts are visible.Backup Browser
The Backups tab on each node includes a file browser that lists backup files from both local storage and the cloud target. Features:
- Scan for Backups — creates a
list_backupsjob to scan local/backups/on the node - Unified file table — shows filename, size, date, and location (Local / Cloud / Both)
- Upload to cloud — offered on rows that exist only on the node, when the node has an enabled cloud target. Creates an
upload_backupjob that pushes that one file from the node to the target. The transfer runs on the node, where the file already is; routing it through the control plane would drag the archive down and push it straight back up. The local copy is kept regardless of the node's delete-after-upload setting — an operator asking for an offsite copy of a file they are looking at did not ask for that file to disappear, and deleting stays an explicit action. The button waits for the job's real verdict, so a failed transfer reports as failed with a link to the job output rather than reading as done - Delete — single Delete button per row that removes the file from every location it exists in (local, cloud, or both); the confirmation dialog names the file and locations explicitly
- Restore Full Project — for
.tar.gzarchives, see therestore_projectrow in the Job Types table - Restore points (incremental chains) — a second table listing each chain on the node's shelf with its runs, size and newest restore point, read from the chain's own
manifest.jsonbyBackupChainListHelper. Restoring picks a run: the full, then every incremental up to it, in order. Chain artifacts are deliberately absent from the flat file table above — listed there,files-0003.tar.gz.encinvites a restore of one incremental with no full under it, which restores nothing at all
What a restore asks, and what it decides
Every restore form asks one thing and decides the rest.
It asks for the domain, pre-filled from the node's recorded URL. This is the one value a restore cannot work out for itself: a rebuild keeps the site's own domain and cuts DNS afterwards, while a rehearsal must not claim it, and the same backup on the same node wants opposite answers. A node provisioned during an incident carries whatever hostname somebody typed in a hurry, so adopting it silently is a mistake that surfaces only after DNS moves.
It decides the serving config. There is no Apache choice on the form. The restore regenerates the virtualhost for this machine from the platform's own templates and never installs the one the backup carries; a differing capture is preserved as {site}.conf.from-backup and named in the job output. On a container node a further step publishes the domain on the host with manage_domain.sh, because the host's proxy virtualhost is outside the container and therefore in no backup.
Every restore job ends with two gates: the site's identity must match the machine (domain, deployment shape, and a database that opens with this machine's credentials) and the site must actually be served — over HTTPS when the domain already resolves here, or reported as certificate-deferred when it does not. The HTTPS gate is explicit because an HTTP-only check passes comfortably while a site serves under a container's internal virtualhost with a valid certificate sitting unused on disk.
What gets reconciled, and why each item is on the list, is in Backups.
Cloud listings are fetched live via TargetLister on every page render (one SigV4 HTTP GET, ~200–500ms). The local listing comes from the most recent completed list_backups job; both the Backups and Database tabs auto-trigger a refresh on page load when that scan is more than 60 seconds stale, so the listing is effectively always current. Both the merge logic and the staleness window are owned by BackupListHelper::get_for_node().
Stored Backups (target-side)
The Backup Targets edit page has a Stored Backups panel that lists the target's objects directly from the bucket and groups them by site. It runs entirely on the control plane via TargetBackups (which lists through S3Signer::list, a continuation-token-paged ListObjectsV2), so it needs no live node — the authoritative view of what is actually stored offsite. Each group is tagged against the node table:
- live — a current node owns the slug; a link jumps to that node's Backups tab for granular local+cloud management
- decommissioned — a soft-deleted node owned the slug; the site is gone but its offsite backups remain here, reachable and deletable
- orphaned — no node, present or deleted, matches the slug
S3Signer from the control plane: a single object (guarded so the key must sit under the target's own prefix), or a whole site's prefix (type-to-confirm the slug). This is the deliberate path for erasing a retired site's offsite backups — deleting a node never touches them.Retiring a node
Two distinct actions on the node detail Overview tab, both permission-10 and CSRF-guarded:
- Remove from Dashboard — soft-deletes the node record only. The site keeps running on its host; Server Manager simply stops tracking it. For a box handed back to its owner or managed elsewhere.
- Permanently Delete Site — creates a
decommission_nodejob that shipsremove_account.shto the host, runs it (-y), and re-probes to confirm the container, its{site}_*volumes, and the reverse-proxy vhost are all gone (DECOMMISSION_VERIFIED). Only on that verification does the result processor soft-delete the node record; a failed or unverified teardown leaves the node intact and enabled to retry. Type-to-confirm the site name; the name is derived from the node's own fields, never operator input. Relays are refused (they tear down through the relay flow).
Removed sites are hidden from the dashboard by default. The Show all sites (including removed) link at the bottom of the Hosts & Sites panel re-renders with them included, each carrying a Removed badge and linking into its still-reachable node detail page (?show_all=1).
Opening a removed node's detail page offers two follow-up actions in its Danger Zone:
- Permanently Delete Site — the same
decommission_nodehost teardown, for a node that was only removed from the dashboard while its site kept running (e.g. an orphaned container). For a removed node it is offered only when this control plane once saw a live site there — a recorded status check, Joinery version, or uptime result. With no such evidence (for example an install that failed and never stood a site up) the action is hidden behind a short note and only Permanently Delete Entry is offered, since there is nothing on the host to tear down. (The page cannot probe the host directly — the web user holds no host SSH key — so this uses evidence already on the record; thedecommission_nodejob itself is idempotent and reportsREMOVE_ACCOUNT_NOTHINGif it reaches a host with nothing to remove.) - Permanently Delete Entry — hard-deletes the Server Manager record itself (
purge_node). Offered only for an already-removed node — purging a still-tracked node is refused, since that is how a live site becomes an untracked orphan. It is also refused while the node's slug still has offsite backups on any enabled target (or while a target cannot be listed to confirm): deleting the record would orphan those backups from the node they belong to, so they must be cleared from the target's Stored Backups panel first. Once allowed, the host is not touched and the job history survives the purge (cascade rules null the references).
Backup Encryption and Key Custody
Default Behavior
Encryption is enabled by default on both Database Backup and Full Project Backup forms. backup_database.sh / backup_project.sh encrypt with AES-256-CBC (PBKDF2, random salt) using the key minted for that run and passed as --key-file. The project archive is encrypted as tar streams into openssl, so the plaintext archive never lands on disk; the artifact is .tar.gz.enc. When a node's backup target is Backblaze B2 encryption is mandatory: the UI replaces the checkbox with a message and the server enforces it regardless of form input.
Key model: one envelope per backup
Every backup run mints its own random encryption key. The archive is encrypted
with it, and the key itself is sealed to two recipients and written beside the
archive as a JSON envelope ({archive}.keys.json), which is uploaded with it:
- recovery — the operator's recovery public key. The private half lives in a password manager and never touches a server. Because a site holds only the public half, the same public key can be configured on any number of sites, so one private key opens every backup from the whole fleet.
- site — a keypair the node itself holds at
config/backup_site_key. This is what lets a site restore itself with nobody present: pre-restore rollback snapshots and routine restores need no operator. It is disposable — lose it and the recovery key still opens everything, and the next run mints a new one.
- The recovery keypair is generated with
maintenance_scripts/sysadmin_tools/escrow_keypair.php(standalone PHP + sodium, no platform bootstrap, so it runs during disaster recovery when the platform is gone), or in the browser from the setup panel. - The public key is stored in the core
backup_recovery_public_keysetting. - Minting and sealing happen on the node, in
maintenance_scripts/sysadmin_tools/backup_envelope.php. Only the recovery public key travels in the job step, so aManagementJobrow — which persists forever — carries nothing that can open anything. - The plaintext key exists only as a 0600 file for the length of the run and is shredded by the step that seals the envelope to the finished archive.
config/backup_site_key is pinned to 600 www-data:www-data by
fix_permissions.sh. A key that exists but cannot be read is an error, never
treated as absent — minting over a live key would orphan the site recipient for
every backup already sealed to the first one.Possession check
Sealing to a public key always appears to succeed, including when the pasted key
is wrong — every backup would then be permanently unopenable, discovered only
during a real recovery. So the key is honored only after the operator unseals a
challenge with the private key. Until that proof is recorded
(backup_recovery_public_key_proven_fpr), BackupRecoveryKey::public_key()
throws and encrypted backups refuse to run.
The check runs against the copy of the key the operator is actually keeping, which is the copy that has to work in a disaster. Two ways to do it, both proving possession of the same X25519 secret:
- In the page — paste the key (from a password manager, typically) into the
setup panel.
BackupRecoveryKey::browser_challenge()packages the proof string asephemeralPub[32] || iv[12] || ciphertext || tag, sealed with X25519 → HKDF-SHA256 (infoBackupRecoveryKey::BROWSER_INFO+ ephemeral public + recipient public) → AES-256-GCM, sobackup_key_verify.jsandassets/js/recovery-readiness.jsopen it with WebCrypto alone. The HKDF context is sent to the browser with the challenge rather than hardcoded at both ends. The key is read from an input outside the form, used in memory, and cleared; it is never submitted, stored, or sent anywhere. Only the recovered proof string is posted, and the server re-checks it. - At the command line —
escrow_keypair.php unsealopens the libsodium sealed-box form of the same challenge with a key file.
Replacing a proven key is a rotation, not an edit: backups already made carry keys sealed to the old public key. Pasting over a proven value is refused.
Guided setup
Recovery key setup is core, not fleet — a standalone site needs it just as much —
so it lives on the Backups page and is rendered by
includes/RecoveryKeySetupPanel.php (see
Backups). The Backup Targets page
shows the current state and links there rather than carrying a second copy of
the panel. BackupRecoveryKey::setup_state() is the single source of truth for
that state, so the panel, the node Backups tab, and the dashboard cannot
disagree.
Until setup is finished the platform does not offer backups it would refuse: a node with a cloud target shows the explanation in place of the Run Backup forms, and a local-only node has encryption switched off with a link to the panel. Creating an encrypting backup checks the recovery key when the job is built, so an unconfigured or unproven key fails while the operator is looking at the button rather than part-way through a backup.
Backups across the fleet
This control plane takes its own backups of the nodes it manages. They are a
separate party's copies of each site, under this control plane's recovery key, on
this control plane's shelf — the manager profile described in
Backups. A site's own
backups are the site profile: its own key, its own schedule, its own business.
Neither owns the other. A site that takes no copies of its own is still backed up from here; a site that takes plenty is still backed up from here. Nothing on either side needs the other to be absent.
The node does the work. backup_run hands it the bucket, a credential and a
recovery public key on stdin, and its own BackupRunner builds the archive,
extends the chain, seals the envelope, uploads and sweeps its local copies.
Routing archives through the control plane would drag every byte down and push it
back up, and would put this machine in the path of every restore.
Nothing is left on the node. The credential is substituted into the step by the agent at run time and never written to a job row or a node's database. The recovery key travels as a public key, which seals but cannot open. So a node holds no key to anyone's backups but its own, and a node that leaves this fleet takes nothing with it.
The node may write to the shelf but never erase it
A backup target holds two credential slots. The main credential
(bkt_credentials) is the control plane's own — it lists, prunes and downloads.
The node credential (bkt_node_credentials, on the target edit form) is an
optional second key created write-only — writeFiles without deleteFiles
on B2, s3:PutObject without s3:DeleteObject on S3. When it is set, that is
the key nodes are handed during a run: a node can add its archives and remove
nothing. When no node credential is configured, nodes receive the main key —
functional, but a compromised node then briefly holds a key that could erase
the shelf, so a fleet target wants the node slot filled.
FleetBackupRetention prunes from here, with the delete-capable main
credential that never leaves this machine. A credential that can delete is a
credential that can erase the fleet's backups, which is the first move of any
ransomware worth the name and the exact thing these copies exist to survive.
Pruning is driven by a bucket listing, which is the opposite of what a site
does for its own backups, and correct only here: this control plane defined the
whole {prefix}/{slug}/manager/ path, knows every slug under it, and is the only
party that can delete from it. It is also stricter — it keeps the newest N sets
of objects that actually exist, so a run that failed part-way can never be
counted as a restore point. Chains are grouped by their directory, so they are
kept or deleted whole by construction.
Two provider notes:
- Linode Object Storage keys are read-only or read-write per bucket with no separate delete capability, so write-without-delete cannot be expressed there. B2 and S3 both express it cleanly.
- A chain rewrites
manifest.jsonevery run. That is a PUT over an existing key, which write-only permits, but on B2 it leaves superseded versions the node cannot remove. Give the fleet bucket a lifecycle rule keeping only the current version.
Scheduling
The Fleet Backups task (plugins/server_manager/tasks/FleetBackupRun.php)
runs every cron tick, finds due nodes, prunes each one's shelf, and dispatches
one backup_run per node.
FleetBackupPolicy resolves each node's schedule: the declared fleet settings,
then that node's own mgn_backup_policy overrides. The fleet default is
enabled. That default is what stops a newly managed node being forgotten —
there is deliberately no detector for "nobody has decided about this node",
because a node nobody decided about is backed up anyway.
The node detail Backups tab edits the policy, as one of three positions:
- Fleet default stores nothing, so the node follows the fleet settings — including future changes to them.
- A schedule of its own stores the full field set (frequency, window, mode, retention, full interval), frozen against the fleet default: a value the operator saw and saved is a value they chose.
- Off stores exactly that decision, which is what lets the dashboard treat a node without fleet backups as somebody's choice rather than a gap.
backup_run the schedule
dispatches, with mode and full-interval taken from the node's policy, so a
manual run extends the same family of restore points the schedule builds.Three rules keep a fleet from behaving like a thundering herd:
- each node's minute is derived from its slug and spread across a window (default 03:00 UTC, 120 minutes wide), so forty nodes do not all begin a multi-hundred-megabyte upload at once;
- a node whose previous run is still pending or running is skipped, so a slow node gets fewer backups rather than a queue;
- no more than
server_manager_fleet_backup_max_concurrentrun at once.
What is reported, and what raises an alarm
check_status asks each node's management API for both profiles: whether each is
scheduled, when it last ran, how it went, whether it reached the bucket, and
which recovery key it sealed to. The manager profile's answer is denormalised
onto mgn_last_backup_time and mgn_last_backup_outcome so the dashboard reads
columns instead of visiting nodes.
The dashboard alarms only on this control plane's own runs —
NodeMonitorHealth::fleet_backup_problems() raises a node whose last backup from
here failed, or whose backups have stopped arriving within its schedule's window.
The alarm is "my backups of this node are broken", not "this node is
unprotected", which is not this control plane's call to make.
The node's word is cross-checked against the bucket. The retention pass
lists each node's shelf with this control plane's own credential before every
run, and the scheduler stamps what it saw — when the shelf was listed and the
newest object write on it — onto mgn_backup_shelf_checked_time and
mgn_backup_shelf_newest_time. The health check compares that against the
node's claimed last run: a shelf listed after a claimed success that holds
nothing written since raises "Backups are not landing". The shelf is the
one witness a compromised or misconfigured node cannot talk into its story —
everything else in the health picture is the node reporting on itself.
A site's own profile is shown alongside as information and never alarms. A node with fleet backups switched off produces nothing either — that was somebody's decision.
Which key each site holds for itself
set_recovery_key.php --report is asked during check_status, and the answer
lands on mgn_backup_recovery_fpr. It prints one machine-readable line,
RECOVERY_KEY=written|proof_write|already|different|none|invalid, and the fleet
table reads the column rather than reaching out to every node on page load.
It is reported and never written. That slot holds the key for the site's own backups; writing it from here would make this control plane the holder of the private half of a key the site believes is its own. There is no job type that can write it. A site whose operator wants a key of their own sets it up on that site's Backups page, with the possession ceremony that makes it trustworthy.
Disaster recovery
To rebuild a lost node from its offsite backups:
- Fetch the archive and its envelope (
{archive}.keys.json) from the bucket. - Recover the archive key on a machine holding the recovery private key:
php backup_envelope.php open --sidecar {archive}.keys.json --private /path/to/recovery.key --key-out /tmp/k - Restore through the dashboard, or with
restore_database.sh --key-file /tmp/k/restore_project.sh --key-file /tmp/k.
restore_project.sh finds the envelope beside
the archive and opens it with config/backup_site_key.This works when the control plane itself is the casualty — the envelopes sit in the bucket alongside the archives, so bucket credentials plus the password-manager private key are sufficient. No site's recovery depends on any other site being alive.
The agent signing key (the fleet trust root) needs no separate recovery record:
it lives at config/agent_signing_key, inside the project tree that the site's own
encrypted project backup carries.
How It Works: Smart Plugin, Dumb Agent
All job-type intelligence lives in JobCommandBuilder.php. The Go agent is a generic executor that understands four primitives: ssh, scp, local, and api.
When an admin triggers an operation:
- PHP looks up the node's connection details (host, SSH key, container, etc.)
JobCommandBuilder::build_<type>()generates an ordered array of steps- PHP writes a job row with the steps in
mjb_commands(JSON) - Go agent picks up the job, executes each main step in order, streams output
- Agent runs the job's teardown steps (if any), then marks the job completed or failed
JobResultProcessoroptionally parses the output into structured data
check_status job looks like in the database:{
"steps": [
{"type": "ssh", "label": "Check disk usage", "cmd": "df -h /"},
{"type": "ssh", "label": "Check memory", "cmd": "free -m"},
{"type": "ssh", "label": "Check uptime", "cmd": "uptime"},
{"type": "ssh", "label": "Check PostgreSQL", "cmd": "pg_isready"},
{"type": "ssh", "label": "Check Joinery version",
"cmd": "grep VERSION /var/www/html/site/public_html/includes/version.php"},
{"type": "ssh", "label": "Container stats",
"cmd": "docker stats --no-stream empoweredhealthtn", "on_host": true}
]
}The agent doesn't know this is a "status check." It just runs each step's command via SSH, captures output, and moves on.
Execution phases: main and teardown
A job's steps form two phases. Steps without the teardown flag are the
continue_on_error) stops
the phase and determines the job's outcome. Steps flagged "teardown": true
are the teardown phase: they run on every exit path — success, mid-job
failure, or none-of-the-main-steps-ran — so the scratch files a job creates
(dumps, staged archives, unpacked installers) are removed even when the job
aborts on a shared production host.Teardown semantics:
- Teardown never changes the outcome. A failed job stays failed with the
original failing step in
mjb_error_message; a teardown step erroring is logged under the=== Teardown ===output header and ignored. - Teardown runs before the terminal status is written. The job stays
runningwhile teardown executes, so the per-node concurrency lock holds (no re-run can race the deletions) and the job detail view keeps streaming. - Progress counts main steps only.
mjb_total_stepsexcludes teardown steps and teardown output never advancesmjb_current_step. - Stale-job replay. Jobs force-failed at agent startup (left
runningby a crash or restart) get their teardown steps replayed frommjb_commands— safe because every teardown command is an idempotentrmon a per-job path. - Placement. Builders put teardown steps at the tail of the array, after
every main step, and keep
continue_on_erroron them. An agent that ignores the flag runs the array sequentially, so tail placement makes the steps plain trailing cleanup there — correct, just not failure-proof.
Adding a New Job Type
Adding a new operation requires PHP changes only -- no Go rebuild needed.
- Add a static method to
JobCommandBuilder:
// plugins/server_manager/includes/JobCommandBuilder.php
public static function build_restart_apache($node) {
return [
['type' => 'ssh', 'label' => 'Restart Apache',
'cmd' => 'systemctl restart apache2'],
['type' => 'ssh', 'label' => 'Verify Apache status',
'cmd' => 'systemctl is-active apache2'],
];
}- Add a UI trigger (button/form) in the appropriate admin view that calls:
$steps = JobCommandBuilder::build_restart_apache($node);
$job = ManagementJob::createJob($node->key, 'restart_apache', $steps, null, $session->get_user_id());
header('Location: /admin/server_manager/job_detail?job_id=' . $job->key);- Optionally add a result processor method in
JobResultProcessorif you want to parse the output into structured data.
Step Fields Reference
| Field | Required | Description |
|---|---|---|
type | Yes | ssh, scp, local, or api |
label | Yes | Human-readable description (shown in UI and output) |
cmd | ssh/local | Shell command to execute |
node_id | No | Override target node (defaults to job's node). Used for multi-node operations like copy_database |
on_host | No | If true, run on the SSH host directly, not inside the Docker container. Used for docker stats, etc. |
direction | scp | upload (local to remote) or download (remote to local) |
remote_path | scp | File path on the remote host |
local_path | scp/api | File path on the control plane (for api, set to stream the response body to a file instead of appending to job output — used by backups/fetch) |
method | api | HTTP method: GET, POST, PUT, DELETE (in practice always GET — the management API is read-only) |
endpoint | api | Path relative to /api/v1/management/ — e.g. stats, backups/list, backups/fetch |
expect_status | api | HTTP status code that counts as success (default 200) |
query | api | Object of query-string params (e.g. {"path": "/backups/foo.sql.gz"}) |
body | api | Request body object (serialized as JSON; ignored for GET/DELETE) |
continue_on_error | No | If true, don't abort the job when this step fails |
timeout | No | Max seconds for this step (default: 1800 = 30 minutes; teardown steps carry 120) |
teardown | No | If true, the step is teardown-phase: it runs on every exit path, its failure never affects the job outcome, and it must be an idempotent removal of a per-job scratch path. Always placed at the tail of the step array with continue_on_error set. |
Management API (Read-Only)
Every Joinery instance exposes a namespaced read-only HTTP surface at /api/v1/management/*. The control plane prefers this over SSH for observability operations (check_status, list_backups) because it's faster, parallelizable, and auditable.
Endpoints (all under /api/v1/management/, all GET, all JSON except backups/fetch which streams binary):
| Endpoint | Replaces SSH step(s) |
|---|---|
health | (new — liveness probe) |
stats | all steps of check_status |
version | Check Joinery version |
databases | List databases |
errors/recent | Recent errors |
backups/list | list_backups |
backups/fetch?path=... | (no control-plane consumer — streams a backup file as binary) |
GET /api/v1/management returns every endpoint with its description.Authentication uses the existing API key system (apk_api_keys — same key headers and hashing as public CRUD; resolved by ApiAuth::authenticate()). The gate (ApiAuth::authorize(), with requires_machine_key + min_user_permission: 10) has two requirements: the key must be a machine key (apk_type = machine) — user session keys minted via /api/v1/auth/login get 403 here, so a superadmin logging into a phone app can't reach the control plane — and its owning user must be a superadmin (usr_permission >= 10). apk_permission is NOT a gate here — it's the CRUD-axis capability and is orthogonal. A superadmin's machine key with apk_permission=1 can call management endpoints; a permission-5 admin's key cannot, regardless of apk_permission.
Adding a management key for a node: on the target node, Admin → API Keys → New Key (admin-created keys are machine keys, which is what the control plane requires), owner = a superadmin user, apk_permission = 1, IP-restrict to the control plane's egress IP. Paste the public/secret pair into the node's Overview tab on the control plane's Server Manager ("API Credential" panel).
> IP restriction on docker-prod nodes: for sites fronted directly by host Apache (no Cloudflare), the container now reads the real client IP via mod_remoteip + the host's X-Forwarded-For: %{REMOTE_ADDR}s header, so IP restriction works end-to-end. For Cloudflare-fronted sites, the container sees Cloudflare's edge IP — IP restriction is not yet meaningful in that case (a future spec will trust Cloudflare's ranges and read CF-Connecting-IP).
Build-time routing: JobCommandBuilder::build_<op>() dispatches to build_<op>_api() or build_<op>_ssh() based on has_api($node, $op), which checks: (1) credentials stored on the node row, (2) a matching build_<op>_api exists, (3) a fresh /health probe succeeds. No runtime fallback — a job is decided at build-time and runs that path or fails. The existing SSH implementation stays in place; clearing the stored credentials or breaking /health routes the next job back to SSH automatically.
Adding a new management endpoint: drop a file under includes/management_api/<name>_handler.php with <name>_handler($request) + <name>_handler_api() meta function. Nested paths mirror directories (backups/list_handler.php → GET /api/v1/management/backups/list). Parallels the action-endpoint convention in logic/*_logic.php. The machine-key + superadmin default applies automatically; a handler can tighten it (never loosen) by returning an 'auth' block from <name>_handler_api() — e.g. 'auth' => ['capability' => 'delete'] for a destructive endpoint. See docs/api.md.
TLS verification is strict by default. The mgn_tls_insecure boolean on mgn_managed_nodes opts a single node out for dev/local instances without a cert from a trusted CA. Audit: SELECT mgn_slug FROM mgn_managed_nodes WHERE mgn_tls_insecure = true.
Data Models
ManagedNode (mgn_managed_nodes)
Represents a remote Joinery instance. Key fields:
mgn_name-- Display name (e.g., "Empowered Health Production")mgn_slug-- Short identifier, unique (e.g., "empoweredhealthtn")mgn_host-- SSH host (IP or hostname)mgn_ssh_user,mgn_ssh_key_path,mgn_ssh_port-- SSH connection detailsmgn_container_name-- Docker container name (null for bare metal)mgn_web_root-- Path topublic_htmlinside the server/containermgn_last_status_data-- JSON from last status check (disk, memory, load, etc.)mgn_joinery_version-- Last known version stringmgn_bkt_backup_target_id-- FK to backup target (null = local only)
CustomerCloudAccount (cca_customer_cloud_accounts)
A user's OAuth-linked cloud provider account (one row per user + provider).
Holds the SecretBox-encrypted token set via storeToken()/getToken();
cca_status is active, refresh_failed, or revoked (the latter two mean
the buyer must re-connect).
CustomerCloudProvision (cvp_customer_cloud_provisions)
One cloud-instance provision, request to running site. cvp_origin is
order (keyed to the getjoinery order item — cvp_external_order_item_id,
unique, required for this origin) or admin (no order item); cvp_status is
the state machine documented under
Customer-Cloud Fulfillment; install parameters
ride on the row (cvp_docker_mode, cvp_install_mode, cvp_source_node_id,
cvp_backup_source, cvp_port, cvp_sitename); links to the account
(cvp_cca_account_id), instance (cvp_instance_id/_ip), and resulting
node (cvp_mgn_node_id).
ManagementJob (mjb_management_jobs)
Represents a queued, running, or completed operation. Key fields:
mjb_mgn_node_id-- Target node (FK to mgn_managed_nodes, null for local-only jobs)mjb_job_type-- Label for display/filtering (e.g., "backup_database")mjb_status--pending,running,completed,failed, orcancelledmjb_commands-- JSON with the step array the agent executesmjb_output-- Progressive text output (appended during execution)mjb_result-- Structured JSON populated byJobResultProcessorafter completionmjb_current_step/mjb_total_steps-- Progress tracking
$job = ManagementJob::createJob(
$node_id, // target node ID (or null for local)
'backup_database', // job type label
$steps, // array of step dicts from JobCommandBuilder
['encryption' => true], // parameters (stored for reference/re-run)
$session->get_user_id() // who triggered it
);BackupTarget (bkt_backup_targets)
Configured storage target for backups. Key fields:
bkt_name-- Display name (e.g., "Production B2")bkt_provider--b2,s3, orlinodebkt_bucket-- Bucket name (required)bkt_path_prefix-- Path prefix within the bucket (default:joinery-backups)bkt_credentials-- JSON with the unified shape{access_key, secret_key, region, endpoint}for every provider; B2's region/endpoint are auto-detected at save timebkt_node_credentials-- optional write-only key handed to nodes during a backup run in place of the main one; same shape, same sealing; B2/S3 onlybkt_delete_local-- Whether to delete local backup after successful uploadbkt_enabled-- Whether this target is active
AgentHeartbeat (ahb_agent_heartbeats)
Single-row table tracking agent liveness. Updated every 30 seconds by the Go agent. The dashboard checks ahb_last_heartbeat to show online/offline status.
Uptime Monitoring
A lightweight per-node uptime check runs on each scheduled-task tick (~15 min). It updates live state on mgn_managed_nodes and emails an admin on up→down and down→up transitions. One alert per transition; no re-alerting while still down.
Augmented mgn_managed_nodes fields:
mgn_uptime_enabled(bool, default true) — per-node on/offmgn_uptime_check_type(varchar, default'http_status') — which check method to use (see below).http_statusis the default because it concludes up/down for any node with a site URL and needs no setup;apiis an opt-in that requires API keys provisioned on the node.mgn_uptime_last_status(varchar) —'up'/'down'/ null (never checked)mgn_uptime_consecutive_failures(int) — streak counter for threshold logicmgn_uptime_down_since(timestamp) — when current outage started, null when upmgn_cert_expiry_ts(timestamp) — observednotAfterof the served TLS cert (see Certificate-expiry monitoring)mgn_cert_alerted_ts(timestamp) — last cert-expiry warning send, for re-alert cadence; null when the cert is comfortably valid
mgn_last_status_check — both check types update it.Per-node IP pinning. http_status checks pin the request to the node's own mgn_host IP (CURLOPT_RESOLVE, SNI/Host preserved) when mgn_host is an IP literal and appears in the site hostname's public A records (DnsResolver::getA()) — the same directly-exposed guard check_cert_expiry() uses. A node behind a shared or round-robin hostname (e.g. two DNS servers sharing dns.scrolldaddy.app via dual A records) is therefore checked as
Check types (extensible via a single dispatch switch in RunNodeUptimeChecks::run_check()):
| Value | Behavior |
|---|---|
api | Reuses JobCommandBuilder::fetch_status_via_api($node). reason='transport' (DNS/connect/timeout) counts as down. 3xx responses also count as down — the API endpoint should never redirect, so a 3xx means the request never reached the API handler (typical cause: infrastructure-level HTTP→HTTPS redirect, possibly looping if Cloudflare is in Flexible mode). Auth (401/403), body errors, and non-3xx non-200 statuses all mean the server responded → up. reason='config' (missing API keys) is a misconfiguration: logged to error log and skipped, no false down alert. |
http_status | Plain curl GET to mgn_site_url. Success = HTTP status in 2xx or 3xx. Forced when mgn_skip_joinery_checks=true regardless of stored check type. |
RunNodeUptimeChecks and a case to the dispatch — no schema change needed.Tick logic (plugins/server_manager/tasks/RunNodeUptimeChecks.php):
- Iterate non-deleted nodes where
mgn_enabledandmgn_uptime_enabledare true andmgn_site_urlis set. - Dispatch on
mgn_uptime_check_type(withmgn_skip_joinery_checksoverriding tohttp_status). - Apply state machine:
- On success: clear failure counter and
down_since, set status'up'. Fire recovered alert on down→up transition. - On failure: incrementconsecutive_failures. Once it reachesFAILURE_THRESHOLD(default 2) and prior status wasn't'down', set status'down', setdown_since=now(), fire down alert. - Inconclusive: the probe records why inmgn_uptime_last_errorand returns without touching status, the failure counter ordown_since, and without alerting.mgn_uptime_last_conclusiveis deliberately left alone, so a node that can never conclude eventually surfaces as stale rather than as healthy.
TIMEOUT_SECONDS=10, FAILURE_THRESHOLD=2. The cron tick interval (~15 min) is the natural rate limiter.A probe only concludes when it reached the node. A failure inside the monitoring host's own name resolution is evidence about the monitoring host, not about the node, so it is inconclusive. Without this, one broken resolver on the control plane fails every probe within a single tick, carries the whole fleet past the failure threshold together, and mails the operator that every site is down while every site is serving traffic — an inverted signal, since the one machine actually at fault is the only one reporting nothing wrong.
NodeMonitorHealth::is_name_resolution_failure($errno, $message) makes the call, and all three check types route through it. It matches curl's CURLE_COULDNT_RESOLVE_HOST/_PROXY by number, and matches on message text for the two cases that carry no distinguishing number: a resolver that hangs rather than answering (curl reports the generic CURLE_OPERATION_TIMEDOUT, wording it "Resolving timed out after…"), and fsockopen, which reports getaddrinfo's text with errno 0. Everything else — refused connections, TLS failures, timeouts once dialling has begun — stays a genuine down result.
The recorded error is worded from the monitoring host's point of view (monitoring host could not resolve <name> (…)) so the dashboard points at the real fault, and the tick's summary line counts these as skipped with the reason attached.
Alert email recipient is resolved per tick via a fallback chain — no new setting:
server_manager_provisioning_admin_alert_email(existing plugin setting)webmaster_email(existing core setting)- The first permission-10 user's email
EmailSender::quickSend() with hard-coded plain-text bodies — no template editor in v1.UI:
- Node edit form (
node_detailoverview tab andnode_add): a "Monitor uptime" checkbox and a "Check type" dropdown. Whenmgn_skip_joinery_checksis on, the runtime forceshttp_statusregardless of the stored value (so pickingapihere is harmless for non-Joinery nodes). - Node detail overview tab: a one-line uptime status under "Last checked" —
Up,Down since X,disabled, ornot yet checked.
Certificate-expiry monitoring
Every enabled node also gets an independent TLS certificate-expiry check on each tick (RunNodeUptimeChecks::check_cert_expiry()), separate from the up/down probe. It warns before a certificate we renew lapses — the failure mode where auto-renewal silently breaks for weeks and the cert expires unnoticed.
It reads the served certificate over the wire (stream_socket_client on ssl://mgn_host:443, capture_peer_cert, SNI = the site hostname), so it sees whatever the cert manager actually serves — Caddy, certbot, anything. Validity is deliberately
notAfter of an already-expired or near-expiry cert is still readable.It self-limits to self-renewed, directly-exposed nodes via two guards, because probing an origin behind a CDN returns a misleading cert:
- Directly-exposed:
mgn_hostmust appear in the public A records for the site hostname (DnsResolver::getA()). A Cloudflare-fronted hostname resolves to Cloudflare, not the origin — those are skipped, because Cloudflare renews that edge cert (not our failure surface). - SAN match: the served cert's CN/SANs must actually cover the hostname. A shared default-vhost fallback cert (SAN mismatch) is ignored — there is nothing dedicated to monitor.
mgn_cert_expiry_ts. When days-remaining drops below server_manager_cert_expiry_warn_days (default 21), it emails a warning through the same recipient fallback chain as the up/down alerts — once on crossing the threshold, then re-alerting every CERT_RECHECK_ALERT_DAYS (default 3) while still under it, and clearing mgn_cert_alerted_ts when a fresh cert pushes the date back out. The node detail overview shows a "TLS cert: expires …" line (warning-styled under threshold) whenever mgn_cert_expiry_ts is set — which also surfaces certs the certbot-file SSL tile can't see (e.g. Caddy nodes).This is distinct from mgn_ssl_state / the SSL tile, which track certbot provisioning status (does an LE cert exist on disk) — a different question from "is the served cert about to expire," and largely a different set of nodes. The two are orthogonal.
Safety Constraints
- Auto-backup before destructive operations --
copy_database,restore_database, andrestore_projectautomatically prepend backup steps.restore_projectsnapshots both the current database (auto_pre_project_restore_*.sql.gz) and the current project tree (auto_pre_project_restore_*.tar.gz) to/backups/before overwriting; either can be skipped if the corresponding component is unchecked in the form. If any pre-backup step fails, the destructive steps never run.
- Database restores replace -- A database restore leaves the target equal to the snapshot. Every restore site (
restore_database, both copy jobs, the from-backup install) verifies the archive withgunzip -tbefore anything is destroyed, drops and recreates thepublicschema so target-only objects are removed too, and loads withpsql -v ON_ERROR_STOP=1so the first load error fails the job instead of completing a partial restore. Dumps are plainpg_dumpsnapshots -- the restore step owns the replacement guarantee, so it holds for any file it is fed. Job-internal dumps (copy jobs, install clone) add--no-owner --no-aclbecause they are restored as the
- Per-node concurrency lock -- The agent skips jobs if another job is already running on the same node, preventing conflicts.
- Stale job recovery -- On agent startup, any orphaned
runningjobs are markedfailedwith a descriptive message.
- Step timeout -- 30-minute default per step, overridable. On timeout, the SSH session is killed.
- Single-threaded agent -- One job at a time. Queued jobs run sequentially.
- Remote credentials at runtime -- Database credentials for backup/copy/restore are extracted from each node's
Globalvars_site.phpat execution time, never stored on the control plane.
API Actions
The dashboard's page JS calls these POST /api/v1/action/server_manager/{name}
actions with the browser-session credential (superadmin only, floor 10) and
reads the response envelope's data. The full set: probe_api, job_status,
discover_nodes, backup_actions, refresh_node_status, add_discovered_nodes.
server_manager/job_status
Polled by the job detail page for live output.
Parameters:
job_id(int) -- job to queryoutput_offset(int) -- character position; only new output since this offset is returned
{
"success": true,
"status": "running",
"new_output": "=== [Step 2/5] Check memory ===\n...",
"output_offset": 1234,
"current_step": 2,
"total_steps": 5,
"error_message": null
}The UI polls every 2 seconds while a job is running and stops when status is completed or failed.
server_manager/backup_actions
Used by the backup browser on the Backups tab.
| Action | Method | Parameters | Returns |
|---|---|---|---|
refresh_list | POST | node_id | {success, job_id} -- creates a list_backups job |
delete_file | POST | node_id, target (local/cloud/both), local_path, cloud_path | {success, job_id} -- creates a delete_backup job |
list_status | POST | node_id, job_id (optional) | {success, status, backup_list, last_scan} -- returns cached file listing |
server_manager/discover_nodes
Used by the auto-detect panel on the Add Node page. Creates and polls discover_nodes jobs.
Troubleshooting
Agent shows Offline on dashboard
- Check the agent is running:
sudo systemctl status joinery-agent - Check logs:
journalctl -u joinery-agent -f - Verify DB credentials in
/etc/joinery-agent/joinery-agent.envmatch those inGlobalvars_site.php
pending forever
- Agent is not running or can't connect to the database
- Another job is running on the same node (per-node lock)
- Verify SSH key path on the node record matches an actual key file
- Test manually:
ssh -i /path/to/key root@host "echo ok" - For container nodes, verify the container name is correct
- The agent crashed or was restarted mid-job. Check
journalctlfor the crash cause. - The partially-completed job should be inspected manually. Use Re-run to retry.
File Reference
Plugin (plugins/server_manager/)
| File | Purpose |
|---|---|
plugin.json | Plugin metadata |
uninstall.php | Removes settings and menu entries on uninstall |
data/managed_node_class.php | ManagedNode + MultiManagedNode |
data/management_job_class.php | ManagementJob + MultiManagementJob |
data/agent_heartbeat_class.php | AgentHeartbeat + MultiAgentHeartbeat |
data/backup_target_class.php | BackupTarget + MultiBackupTarget |
includes/JobCommandBuilder.php | Command generation for all job types; ssh_prefix() is public for use by other tools |
node_exec.php | Dev/AI diagnostic CLI — run any command on a managed node in one call; handles SSH + Docker transparently. Usage: php node_exec.php (list nodes), php node_exec.php <slug> "<cmd>" (run), or pipe stdin with --stdin for SQL queries. |
includes/JobResultProcessor.php | Parses completed job output into structured data |
includes/S3Signer.php | AWS SigV4 signer for S3-compatible storage (get/put/delete) |
includes/TargetUploader.php | Web-tier upload + delete helpers using S3Signer |
includes/TargetLister.php | Web-tier paginated bucket listing using S3Signer |
includes/TargetTester.php | Connection test on Save for Backup Targets |
includes/node_uploader.php | Self-contained upload/delete/download dispatcher run on the node via heredoc; composed at job-build time with S3Signer + injected credentials |
includes/BackupListHelper.php | Merges latest local list_backups job output with live cloud listing into a unified file table |
ajax/job_status.php | Live job output polling |
ajax/discover_nodes.php | Creates and polls node discovery jobs |
ajax/backup_actions.php | Backup browser actions (scan, delete) |
migrations/migrations.php | Indexes, admin menu entries, menu consolidation |
views/admin/index.php | Dashboard -- fleet overview, publish upgrade |
views/admin/node_detail.php | Node detail -- tabbed page (overview/backups/database/updates/jobs) |
views/admin/node_add.php | Add node -- auto-detect + manual form |
views/admin/targets.php | Backup target CRUD |
views/admin/jobs.php | Global job history |
views/admin/job_detail.php | Single job output with live polling |
views/admin/nodes_edit.php | Redirect stub (-> node_detail or node_add) |
views/admin/nodes.php | Redirect stub (-> dashboard) |
views/admin/backups.php | Redirect stub (-> dashboard or node_detail) |
views/admin/database.php | Redirect stub (-> dashboard or node_detail) |
views/admin/updates.php | Redirect stub (-> dashboard or node_detail) |
Go Agent (/home/user1/joinery-agent/)
| File | Purpose |
|---|---|
main.go | Entry point, signal handling, poll loop |
config.go | Environment-based configuration |
db.go | PostgreSQL: job claiming, output writing, heartbeat |
runner.go | Step executor dispatching to ssh/scp/local |
ssh.go | SSH connection pooling and command execution |
scp.go | SCP file transfer |
server.go | Node connection info struct |
Makefile | build, test, release targets |
build_installer.sh | Generates self-extracting installer |
install/joinery-agent.service | systemd unit file |
config/joinery-agent.env.example | Example configuration |