System audit · read-only

What's out of date and what's exposed
on srv1738752

Everything found on the production VPS: the patch backlog that needs a maintenance window, and the security gaps that don't. Nothing was changed, restarted, or written during this audit.

Host srv1738752 · Ubuntu 24.04.4 LTS Kernel 6.8.0-139 (running) Uptime 15 days Audited 02 Oct 2026 Method read-only inspection

Verdict

50
ports answering from the public internet
5
critical findings
56
packages pending update
1
reboot required (kernel installed, not running)
0
security patches outstanding
7
backup sets verified offsite

Nothing is currently broken, and the data is safe. The eval portal, Unified Platform, and every routed site are up. Backups run daily and push offsite encrypted.

The problem is that the box is wide open at the network layer. A firewall that would have blocked it is switched off, and 48 application ports — including the database, a code editor and a writable terminal — are reachable directly from the internet without any password. This is the part that does not need a maintenance window to fix.

The patch backlog is real but ordinary: 56 packages, none of them security-flagged, and one reboot that has been pending since mid-September.

Exposure — what the internet can reach

Ports 80 and 443 are Traefik and are meant to be public. Everything else on this list is published on 0.0.0.0 and answers a connection from outside the VPS, bypassing Traefik's TLS and any login it would have applied.

PortServiceWhy it matters
8081code-serverauth: none — full VS Code with a terminal, no password
7681ttyd--writable web terminal, no credential
5432PostgreSQL 17.6All departmental data (unified, eval, reimbursement)
6543Supabase poolerSecond path to the same database
8002Supabase StudioDatabase admin UI
8001 / 8444Supabase KongAPI gateway and admin API
8003ImmichPhoto library
8787Hermes WebUIThis agent's own control surface
8098Unified Platform + eval portalPlain HTTP, no TLS, direct — no login wall at the network edge
8092 / 8086Agentic OS / data serviceInternal dashboards
8080 / 8095Listening Room API + webInternal app
8005omnirouteLLM gateway
8765tetrisStatic server
8104Master B staticStatic file server
8101, 8105, 8138–8160~25 report/static serversAll plain HTTP, no auth
5098 / 5099FormForge / PDF serviceDocument services
8082TraefikSecond Traefik listener

Bound to loopback only and therefore not exposed: Open Notebook (5055, 8502, 8008), Repo Vault (8103), reimbursement, ink-and-ember, interview-campaign, and the Hermes agent port 32768.

Critical — fix without a maintenance window

Critical

code-server is open to the world with no password

auth: none in the config, bound to 0.0.0.0:8080, published as port 8081. Anyone who finds it gets a browser IDE with a built-in terminal — that is remote code execution on the host, no exploit required.

Fix: set auth: password with a hashed password, or stop publishing the port and reach it only over Tailscale. Neither restarts the Unified Platform.

Critical

ttyd gives away a writable shell

Running as ttyd --port 7681 --interface 0.0.0.0 --writable --cwd /workspace with no credentials. Same class of problem as code-server: an anonymous interactive shell.

Fix: add a credential, bind it to loopback, or stop the service — the terminal is reachable through other doors you already control.

Critical

PostgreSQL is listening on the public internet

The database container publishes 5432 and its pg_hba.conf ends with host all all 0.0.0.0/0 scram-sha-256 — password authentication from anywhere. The pooler on 6543 is a second door to the same data. Passwords are the only thing standing between the internet and every record in unified.

Fix: stop publishing 5432/6543 and keep database access on the Docker network plus the existing SSH tunnel. This is a one-line compose change, not a data change.

Critical

No firewall, no fail2ban

ufw is inactive. Nothing throttles the 50 open ports above, and nothing blocks repeated login failures. Over 30 days there were 333 failed SSH auth attempts; the busiest single source (221.225.89.175) made 205 tries in the last week.

Fix: enable ufw allowing 22/80/443 only, then work down the published ports. Install fail2ban for SSH. Rule changes apply live and affect nothing running.

Critical

SSH accepts root login by password

PermitRootLogin yes and PasswordAuthentication yes, port 22 open to the internet, 6 attempts per connection and no lockout. The host key already carries 7 authorised keys, so password login is not needed by anyone.

Fix: set PermitRootLogin prohibit-password and PasswordAuthentication no, then reload sshd. Existing key sessions stay connected — verify with a second session before closing the first.

Patch backlog — needs a window

56 packages pending, none security-flagged. Automatic security updates are enabled and running — they applied openssl/libssl3t64 on 1 Oct and libxpm4 on 2 Oct. What's left is functional and maintenance updates.

Reboot required

The kernel 6.8.0-142 is installed but the host is still running 6.8.0-139; the system has been flagging *** System restart required *** since about 17 September. A third kernel, 6.8.0-146, is also queued in the pending list, so one reboot clears both.

Notable pending packages

PackageInstalledAvailableNote
linux-image6.8.0-1426.8.0-146running 6.8.0-139 — reboot needed
docker-ce29.4.329.8.2restarts the Docker daemon
containerd.io2.2.32.3.6container runtime
docker-compose-plugin5.1.35.6.0used by the Traefik stack
docker-buildx-plugin0.33.00.37.1
nodejs22.23.122.23.3Node apps need a restart after
tailscale1.102.21.102.4remote access path
apparmor + libapparmor14.0.14.0.1-8security confinement
krb5 (5 pkgs)1.20.1-6.71.20.1-6.10
netplan.io (4 pkgs)1.1.2-8.21.1.2-8.3network config tooling
apport (4 pkgs)2.28.12.28.3
google-chrome-stable154.0.8037.57154.0.8037.97used by headless browser work
xvfb + xserver-common21.1.12-1.621.1.12-1.8
snapd2.762.76.3
base-files, procps, iproute2, dmidecode, plymouth, byobu, open-vm-tools, qemu-guest-agent, multipath-tools, kpartx, sosreport, libxmlb2, libjcat1, libpciaccess0, console-setup, motd-newsroutine maintenance revisions

Housekeeping — worth fixing while you're in there

Medium

Disk is 88% full

169 GB used of 193 GB — 25 GB free. Docker accounts for 35.4 GB of images, 3.9 GB of containers and 78 GB of volumes across 17 volumes. Nothing is at risk today, but a large build or an unbounded log can fill the remainder quickly.

Fix: audit the inactive volumes and the 1.5 GB of build cache. No downtime.

Medium

No swap on a box running at 87% memory

15 GB RAM, 13 GB in use, 2.2 GB available, and Swap: 0B. Under a memory spike there is no cushion — the kernel will pick a process to kill. The largest single consumer is a python server.py at 2.5 GB.

Fix: add a swap file or raise the instance size. Takes effect immediately.

Medium

Nine report servers won't survive the reboot

Nine static servers are running as bare processes with no systemd unit. They will not come back after a restart, and their URLs will return 502 until someone restarts them by hand. The other 14 static servers are properly supervised.

PortContent root
8105/var/www/forex
8152/var/www/callout-line
8154/var/www/abu-letters
8155/var/www/rating-scale
8156/var/www/promotion-audit
8157/var/www/role-dashboards
8158/var/www/walkthrough
8159/var/www/unified-tour (basic auth)
8160no directory argument

Fix: convert each to a systemd unit with Restart=always before the reboot. This is also the one item that makes a reboot risky rather than routine.

Medium

Container images are years behind

Several pinned images have not moved since they were first pulled — kong:3.9.1 (16 months), postgrest:v14.12 (14 months), the Immich Postgres image (12 months), imgproxy:v3.30.1 (11 months). These sit on the data path and have accumulated upstream fixes nobody has pulled.

Fix: upgrade the Supabase stack images on their own window, with a database dump taken immediately beforehand.

Low

An internal process keeps guessing a wrong SSH username

Something on the Hermes container connects to the VPS as user hermeswebui and fails — 31 times in the last week. It is not an attack (the source is the container's own bridge IP), but it is noise that hides real attempts in the log.

Fix: correct the username in whatever automation script uses it.

Already healthy — no action

Verified

Backups are running and verified

A full database dump lands nightly at 03:00 (latest 2 Oct, 3.7 MB), the Unified Platform backup runs at 03:05, and an encrypted offsite copy is pushed and verified at 05:30 — the log confirms 7 sets in the remote repository, newest being a 5.0 MB .gpg archive. Interview season and M4B backups also run on schedule.

Verified

Automatic security updates are on

unattended-upgrades is enabled and active, and it is working — it applied OpenSSL updates on 1 Oct without anyone touching the box. This is why the outstanding count contains zero security patches.

Verified

TLS certificates are current

Every routed host has a valid Let's Encrypt certificate: unified and portal to 28 Nov, scl/schedule/qgenda to 21 Nov, code to 20 Nov, photos to 22 Dec. Traefik renews automatically.

Verified

Data is intact

The evaluation tracker still holds its cycle (AY 2026–2027) with all 57 assignments assigned. The Unified Platform container has 0 restarts and is passing its health check.

Pick a window

Three separate jobs with very different risk. They do not have to happen at the same time — and the first one shouldn't wait for the others.

1 · Close the doors

~30 min · no downtime

Firewall, SSH hardening, stop publishing the database, lock down code-server and ttyd. Applies live. Nothing restarts; the eval portal and Unified Platform keep serving throughout. Can be done any time, including during the workday.

  • Risk: low — the only way to lock yourself out is a firewall rule that drops port 22, so it is added before the policy is enforced, with a second session open
  • Reversible: yes, each change is a one-line revert

2 · Patch and reboot

~15–25 min downtime

Install the 56 packages, clear the kernel reboot, bring up Docker 29.8 and containerd 2.3. Every container restarts; the 14 supervised static servers come back on their own. Needs the pre-reboot checklist below done first.

  • What goes down: all web apps, the Unified Platform and eval portal, Hermes itself, for roughly 10 minutes after the reboot while containers start
  • What breaks if you skip the checklist: nine report URLs return 502
  • Best time: a weekday evening or early Saturday — away from clinic hours

3 · Refresh the container images

~30–45 min · partial downtime

Pull current Supabase, Kong, PostgREST and Immich images and recreate the stack. Do this as its own job — never in the same window as the OS patch, so that if something misbehaves you know which change caused it.

  • Window to avoid: 03:00–05:35, when the nightly database and offsite backups run
  • Prerequisite: take a manual dump immediately before, even though the 03:00 one exists

Suggested order and slots

WhenJobWhy then
Anytime this week1 · Close the doorsNo downtime, highest payoff
Evening, 19:00–21:00Pre-reboot checklist + 2 · Patch and rebootClear of clinic hours and clear of the 03:00 backup run
Following weekend3 · Refresh container imagesSeparate change, separate blame if anything moves

Pre-reboot checklist

Do these before the reboot window, in this order. The first item is the one that actually matters.

Rollback

Firewall and SSH changes revert with the line that made them. For the patch job, the previous kernel (6.8.0-139) stays installed on disk, so a boot into the older kernel remains available from the provider console if 6.8.0-146 misbehaves. Docker packages can be reverted to the version listed in the patch table above.