Verdict
Nothing is currently broken, and the data is safe. The eval portal, Unified Platform, and every routed site are up. Backups run daily and push offsite encrypted.
The problem is that the box is wide open at the network layer. A firewall that would have blocked it is switched off, and 48 application ports — including the database, a code editor and a writable terminal — are reachable directly from the internet without any password. This is the part that does not need a maintenance window to fix.
The patch backlog is real but ordinary: 56 packages, none of them security-flagged, and one reboot that has been pending since mid-September.
Exposure — what the internet can reach
Ports 80 and 443 are Traefik and are meant to be public. Everything else on
this list is published on 0.0.0.0 and answers a connection from outside the VPS,
bypassing Traefik's TLS and any login it would have applied.
| Port | Service | Why it matters |
|---|---|---|
| 8081 | code-server | auth: none — full VS Code with a terminal, no password |
| 7681 | ttyd | --writable web terminal, no credential |
| 5432 | PostgreSQL 17.6 | All departmental data (unified, eval, reimbursement) |
| 6543 | Supabase pooler | Second path to the same database |
| 8002 | Supabase Studio | Database admin UI |
| 8001 / 8444 | Supabase Kong | API gateway and admin API |
| 8003 | Immich | Photo library |
| 8787 | Hermes WebUI | This agent's own control surface |
| 8098 | Unified Platform + eval portal | Plain HTTP, no TLS, direct — no login wall at the network edge |
| 8092 / 8086 | Agentic OS / data service | Internal dashboards |
| 8080 / 8095 | Listening Room API + web | Internal app |
| 8005 | omniroute | LLM gateway |
| 8765 | tetris | Static server |
| 8104 | Master B static | Static file server |
| 8101, 8105, 8138–8160 | ~25 report/static servers | All plain HTTP, no auth |
| 5098 / 5099 | FormForge / PDF service | Document services |
| 8082 | Traefik | Second Traefik listener |
Bound to loopback only and therefore not exposed: Open Notebook (5055, 8502, 8008), Repo Vault (8103), reimbursement, ink-and-ember, interview-campaign, and the Hermes agent port 32768.
Critical — fix without a maintenance window
code-server is open to the world with no password
auth: none in the config, bound to 0.0.0.0:8080, published as
port 8081. Anyone who finds it gets a browser IDE with a built-in terminal — that is remote
code execution on the host, no exploit required.
Fix: set auth: password with a hashed password, or stop
publishing the port and reach it only over Tailscale. Neither restarts the Unified Platform.
ttyd gives away a writable shell
Running as ttyd --port 7681 --interface 0.0.0.0 --writable --cwd /workspace
with no credentials. Same class of problem as code-server: an anonymous interactive shell.
Fix: add a credential, bind it to loopback, or stop the service — the terminal is reachable through other doors you already control.
PostgreSQL is listening on the public internet
The database container publishes 5432 and its pg_hba.conf ends
with host all all 0.0.0.0/0 scram-sha-256 — password authentication from
anywhere. The pooler on 6543 is a second door to the same data. Passwords are
the only thing standing between the internet and every record in unified.
Fix: stop publishing 5432/6543 and keep database access on the Docker network plus the existing SSH tunnel. This is a one-line compose change, not a data change.
No firewall, no fail2ban
ufw is inactive. Nothing throttles the 50 open ports above, and nothing
blocks repeated login failures. Over 30 days there were 333 failed SSH auth attempts;
the busiest single source (221.225.89.175) made 205 tries in the last week.
Fix: enable ufw allowing 22/80/443 only, then work down the published ports. Install fail2ban for SSH. Rule changes apply live and affect nothing running.
SSH accepts root login by password
PermitRootLogin yes and PasswordAuthentication yes, port 22 open
to the internet, 6 attempts per connection and no lockout. The host key already carries 7
authorised keys, so password login is not needed by anyone.
Fix: set PermitRootLogin prohibit-password and
PasswordAuthentication no, then reload sshd. Existing key sessions stay
connected — verify with a second session before closing the first.
Patch backlog — needs a window
56 packages pending, none security-flagged. Automatic security updates
are enabled and running — they applied openssl/libssl3t64 on 1 Oct and
libxpm4 on 2 Oct. What's left is functional and maintenance updates.
Reboot required
The kernel 6.8.0-142 is installed but the host is still running
6.8.0-139; the system has been flagging *** System restart required ***
since about 17 September. A third kernel, 6.8.0-146, is also queued in the
pending list, so one reboot clears both.
Notable pending packages
| Package | Installed | Available | Note |
|---|---|---|---|
| linux-image | 6.8.0-142 | 6.8.0-146 | running 6.8.0-139 — reboot needed |
| docker-ce | 29.4.3 | 29.8.2 | restarts the Docker daemon |
| containerd.io | 2.2.3 | 2.3.6 | container runtime |
| docker-compose-plugin | 5.1.3 | 5.6.0 | used by the Traefik stack |
| docker-buildx-plugin | 0.33.0 | 0.37.1 | |
| nodejs | 22.23.1 | 22.23.3 | Node apps need a restart after |
| tailscale | 1.102.2 | 1.102.4 | remote access path |
| apparmor + libapparmor1 | 4.0.1 | 4.0.1-8 | security confinement |
| krb5 (5 pkgs) | 1.20.1-6.7 | 1.20.1-6.10 | |
| netplan.io (4 pkgs) | 1.1.2-8.2 | 1.1.2-8.3 | network config tooling |
| apport (4 pkgs) | 2.28.1 | 2.28.3 | |
| google-chrome-stable | 154.0.8037.57 | 154.0.8037.97 | used by headless browser work |
| xvfb + xserver-common | 21.1.12-1.6 | 21.1.12-1.8 | |
| snapd | 2.76 | 2.76.3 | |
| base-files, procps, iproute2, dmidecode, plymouth, byobu, open-vm-tools, qemu-guest-agent, multipath-tools, kpartx, sosreport, libxmlb2, libjcat1, libpciaccess0, console-setup, motd-news | routine maintenance revisions | ||
Housekeeping — worth fixing while you're in there
Disk is 88% full
169 GB used of 193 GB — 25 GB free. Docker accounts for 35.4 GB of images, 3.9 GB of containers and 78 GB of volumes across 17 volumes. Nothing is at risk today, but a large build or an unbounded log can fill the remainder quickly.
Fix: audit the inactive volumes and the 1.5 GB of build cache. No downtime.
No swap on a box running at 87% memory
15 GB RAM, 13 GB in use, 2.2 GB available, and Swap: 0B. Under a memory
spike there is no cushion — the kernel will pick a process to kill. The largest single
consumer is a python server.py at 2.5 GB.
Fix: add a swap file or raise the instance size. Takes effect immediately.
Nine report servers won't survive the reboot
Nine static servers are running as bare processes with no systemd unit. They will not come back after a restart, and their URLs will return 502 until someone restarts them by hand. The other 14 static servers are properly supervised.
| Port | Content root |
|---|---|
| 8105 | /var/www/forex |
| 8152 | /var/www/callout-line |
| 8154 | /var/www/abu-letters |
| 8155 | /var/www/rating-scale |
| 8156 | /var/www/promotion-audit |
| 8157 | /var/www/role-dashboards |
| 8158 | /var/www/walkthrough |
| 8159 | /var/www/unified-tour (basic auth) |
| 8160 | no directory argument |
Fix: convert each to a systemd unit with Restart=always
before the reboot. This is also the one item that makes a reboot risky rather than routine.
Container images are years behind
Several pinned images have not moved since they were first pulled — kong:3.9.1
(16 months), postgrest:v14.12 (14 months), the Immich Postgres image (12 months),
imgproxy:v3.30.1 (11 months). These sit on the data path and have accumulated
upstream fixes nobody has pulled.
Fix: upgrade the Supabase stack images on their own window, with a database dump taken immediately beforehand.
An internal process keeps guessing a wrong SSH username
Something on the Hermes container connects to the VPS as user hermeswebui and
fails — 31 times in the last week. It is not an attack (the source is the container's own
bridge IP), but it is noise that hides real attempts in the log.
Fix: correct the username in whatever automation script uses it.
Already healthy — no action
Backups are running and verified
A full database dump lands nightly at 03:00 (latest 2 Oct, 3.7 MB), the Unified Platform
backup runs at 03:05, and an encrypted offsite copy is pushed and verified at 05:30 —
the log confirms 7 sets in the remote repository, newest being a 5.0 MB
.gpg archive. Interview season and M4B backups also run on schedule.
Automatic security updates are on
unattended-upgrades is enabled and active, and it is working — it applied
OpenSSL updates on 1 Oct without anyone touching the box. This is why the outstanding count
contains zero security patches.
TLS certificates are current
Every routed host has a valid Let's Encrypt certificate: unified and
portal to 28 Nov, scl/schedule/qgenda to
21 Nov, code to 20 Nov, photos to 22 Dec. Traefik renews
automatically.
Data is intact
The evaluation tracker still holds its cycle (AY 2026–2027) with all 57 assignments
assigned. The Unified Platform container has 0 restarts and is passing its
health check.
Pick a window
Three separate jobs with very different risk. They do not have to happen at the same time — and the first one shouldn't wait for the others.
1 · Close the doors
~30 min · no downtimeFirewall, SSH hardening, stop publishing the database, lock down code-server and ttyd. Applies live. Nothing restarts; the eval portal and Unified Platform keep serving throughout. Can be done any time, including during the workday.
- Risk: low — the only way to lock yourself out is a firewall rule that drops port 22, so it is added before the policy is enforced, with a second session open
- Reversible: yes, each change is a one-line revert
2 · Patch and reboot
~15–25 min downtimeInstall the 56 packages, clear the kernel reboot, bring up Docker 29.8 and containerd 2.3. Every container restarts; the 14 supervised static servers come back on their own. Needs the pre-reboot checklist below done first.
- What goes down: all web apps, the Unified Platform and eval portal, Hermes itself, for roughly 10 minutes after the reboot while containers start
- What breaks if you skip the checklist: nine report URLs return 502
- Best time: a weekday evening or early Saturday — away from clinic hours
3 · Refresh the container images
~30–45 min · partial downtimePull current Supabase, Kong, PostgREST and Immich images and recreate the stack. Do this as its own job — never in the same window as the OS patch, so that if something misbehaves you know which change caused it.
- Window to avoid: 03:00–05:35, when the nightly database and offsite backups run
- Prerequisite: take a manual dump immediately before, even though the 03:00 one exists
Suggested order and slots
| When | Job | Why then |
|---|---|---|
| Anytime this week | 1 · Close the doors | No downtime, highest payoff |
| Evening, 19:00–21:00 | Pre-reboot checklist + 2 · Patch and reboot | Clear of clinic hours and clear of the 03:00 backup run |
| Following weekend | 3 · Refresh container images | Separate change, separate blame if anything moves |
Pre-reboot checklist
Do these before the reboot window, in this order. The first item is the one that actually matters.
- Convert the nine unsupervised static servers (ports 8105, 8152, 8154–8160) to systemd
units with
Restart=always, so their URLs come back on their own. - Take a manual database dump, even though the 03:00 one exists.
- Confirm the offsite backup push completed that morning — the log should read
OK: pushed. - Note the running container list and their restart policies, so you can compare after.
- Keep one SSH session open throughout the reboot; reconnect on a fresh one to prove the host is genuinely back.
- After the reboot, verify in this order: SSH reachable → Docker daemon up → database container healthy → Unified Platform returning 200 → the eval tracker loading → the nine converted static URLs returning 200.
Rollback
Firewall and SSH changes revert with the line that made them. For the patch job, the previous kernel (6.8.0-139) stays installed on disk, so a boot into the older kernel remains available from the provider console if 6.8.0-146 misbehaves. Docker packages can be reverted to the version listed in the patch table above.