Skip to main content
What changes between releases and in what order to apply it.

The shape of an upgrade

Four things happen, and the order is not interchangeable:
  1. New code arrives (new image, or new checkout plus composer install)
  2. Landlord migrations run — platform-wide tables, including any that the SaaS operator package adds
  3. Tenant migrations run — once per tenant schema
  4. Queue workers restart, so they stop running the previous release from memory
Step 4 is the one most often forgotten. Horizon workers are long-lived PHP processes: until they are restarted they keep executing the old code against the new database schema, which is the worst combination of the two.

Docker Compose

horizon:terminate lets in-flight jobs finish and then exits; Compose restarts the container with the new image. Killing workers outright drops whatever they were processing — and for an inbound webhook that means a message the provider already considers delivered. Pin a version rather than tracking latest if you want upgrades to be a decision rather than an event:

Two schemas, two migration commands

Tenant isolation is schema-per-tenant, so a migration touching tenant tables must run once per tenant. ops:tenants-migrate iterates active tenants and runs each tenant’s schema migrations and then its settings migrations; plain artisan migrate does not and will leave every tenant schema untouched. A release can contain either kind or both. Running both commands unconditionally is safe — an already-applied migration is skipped.

Before upgrading

Back up the database. Migrations are not reversible in practice: down() methods exist but are exercised far less than up(), and a partial rollback on a schema-per-tenant layout is worse than a restore.
Read the release notes for breaking changes, particularly any that touch Flow node contracts. A node handler’s contract is versioned: existing flow definitions keep running against the old version, but a release may add a new one that new flows use.

Moving to host mode

Switching TENANCY_RESOLUTION from single to host changes which hosts the application answers, so do it deliberately:
  1. Set up the wildcard DNS record and certificate first (Several tenants).
  2. Make sure SESSION_DRIVER is not database and SESSION_DOMAIN is unset. Sessions stored by the old driver are not carried over, so everyone signs in again.
  3. Put WEBHOOK_BASE_URL and the gateway host under TENANCY_BASE_DOMAIN. A host outside it gets 400, and channels registered against such a host stop receiving.
  4. Set TENANCY_RESOLUTION=host, then recreate the containers so the route cache is rebuilt: docker compose up -d --force-recreate app horizon scheduler.
The existing tenant keeps its slug and is served at <slug>.<base domain>; the base domain stops serving the panel. Going back to single is the same edit the other way.

Trusted proxies

TRUSTED_PROXIES is new. Compose installations get a default that trusts only the bundled Caddy at its fixed address on the compose network, which makes the client address correct behind it and changes what rate limits see: they now count real clients instead of one. The Compose file now also defines the network’s subnet (172.30.0.0/24, dynamic addresses from 172.30.0.128/25); set COMPOSE_SUBNET, COMPOSE_IP_RANGE and CADDY_ADDRESS together if that collides with something on your host, and give a second installation from the same file (a staging ENV_FILE) its own values. Because the network definition changed, recreate it once: docker compose down and then docker compose up -d. Bare metal and other setups trust nobody until the variable is set, as before.

Host-mode installs: sessions are now checked

An installation already running with TENANCY_RESOLUTION=host that has SESSION_DRIVER=database or SESSION_DOMAIN set used to fail only on platform pages. After this upgrade every web request fails with a message naming the setting. Before upgrading, switch to SESSION_DRIVER=redis and unset SESSION_DOMAIN; people sign in again once. Queue workers and console commands are not affected.

Flow calls to internal addresses are refused

The call node no longer connects to private, loopback, link-local, reserved or cloud-metadata addresses, and this applies to every installation. A flow that calls an internal service (http://crm.local, http://10.0.0.7/api, a mock on localhost) now takes its error handle with error_code = egress_denied instead of getting a response. Redirects are checked on each hop, so a public URL that redirects to an internal one is refused too. Before upgrading, list the internal targets your flows rely on in FLOW_EGRESS_ALLOW (for example FLOW_EGRESS_ALLOW=10.20.0.0/16,crm.internal), then restart Horizon. Cloud metadata addresses cannot be listed. Calls also stop honouring the HTTP_PROXY and HTTPS_PROXY environment variables. If your calls leave the network through a proxy, set FLOW_EGRESS_PROXY to the same address; that proxy must refuse private ranges itself. Names from /etc/hosts or container aliases still resolve, but their addresses are checked like any other, so allow them by range or name (a mock on localhost needs FLOW_EGRESS_ALLOW=127.0.0.1/32). See Configuration.

Downtime

The application tolerates a short window where old and new run side by side — webhook ingress answers, and work queues up rather than being lost. What it does not tolerate is old workers against a migrated schema, which is why they restart last. For a zero-downtime upgrade the ordering is: start new web containers, migrate, then restart workers. Anything more elaborate is untested territory here.

The gateway

The gateway upgrades independently of the application — it shares only Redis and the queue format, both versioned.
If a release changes the queue payload version, the gateway and the application must be upgraded together: the worker rejects a payload version it does not recognise, sending the job to failed_jobs rather than processing it with half its context. Release notes call this out when it applies.

Rolling back

Code rolls back cleanly. The database does not. If the release you are leaving contained migrations, restore the backup instead — or confirm from the release notes that its migrations are additive, in which case older code simply ignores the new columns.

Verifying