Skip to main content

Patching and system setup

Prepare and patch servers before application deployment.

Ansible hosts group: patchservers​

Vault data source​

All secrets are fetched at runtime from the Vault KV service.

  • Create a bucket named after cs_project_code.
  • Secrets are under {{ cs_project_code }}/application-deployer/clusters.

DevOps server​

  • key: {{ cs_project_code }}/devops-server/hosts/{{ DEVOPS_SERVER_INVENTORY_HOSTNAME }}/apps/scm/generated

DEVOPS_SERVER_INVENTORY_HOSTNAME is a mandatory environment variable holding the DevOps server host's inventory hostname. The artifact registry URL and admin API token that cs_vm_artifact_registry_* use come from this key; it is written by devops-server-deployer.

Networks​

  • key: {{ cs_project_code }}/application-deployer/clusters/{{ Lab Name or Cluster name }}/networks/main
{
"cidr": "IPv4 CIDR covering all hosts in the cluster network.",
"dns": ["List of DNS servers for the cluster."],
"vpn_cidr": "IPv4 CIDR for the VPN network (e.g. WireGuard)."
}

Hosts​

  • key: {{ cs_project_code }}/application-deployer/clusters/{{ Lab Name or Cluster name }}/hosts/{{ hostname }}
{
"service_password": "Service account static password",
"host": "Host IPv4 address",
"root_password": "root account password",
"dockerd_tcp_port": "(int) Docker sock TCP/TLS port.",
"glances_monitoring_password": "Password protecting the Glances web server / REST API",
"public_keys": "List of the host's own SSH public keys, written as known_hosts entries for `host`"
}

Certificate authority​

  • key: {{ cs_project_code }}/certificate-authority/root-ca
{
"root_ca_cert_pem": "PEM-encoded root CA certificate.",
"root_ca_key_pem": "PEM-encoded root CA private key.",
"root_ca_key_password": "Passphrase for the root CA private key.",
"domain": "Root CA domain name (used to derive the cluster's domain name)."
}

Used by all services that issue or verify TLS certificates (reverse proxy, PostgreSQL, Docker API, MinIO, Emby, etc.). The playbook signs leaf certificates with the root CA private key using the community.crypto ownca provider and distributes the root CA certificate as the trust anchor. cs_force_rotate_leaf_certs forces every private key, CSR, and certificate through its own regeneration path: a private key sets regenerate: always, a CSR or certificate sets force: true. See Global variables for cs_force_rotate_leaf_certs.

Managed services​

  • key: managed-services/smtp (its own managed-services bucket, outside the {{ cs_project_code }} bucket)
{
"smtphost": "SMTP server hostname.",
"smtpport": "(int) SMTP server port.",
"smtpusername": "SMTP auth username.",
"smtppassword": "SMTP auth password.",
"fromaddress": "Envelope/From address used for outbound mail."
}

A single shared SMTP endpoint used by every service that sends outbound email (Vikunja, Nextcloud, qBittorrent Nox, WireGuard peer-notification mail, etc.).

LUKS disks​

  • key: {{ cs_project_code }}/application-deployer/clusters/{{ Lab Name or Cluster name }}/hosts/{{ hostname }}/disks/{{ disk name mount point }}

Disk must be LUKS Encrypted. And the data partition must be formatted as ext4.

Get crypt_device_uuid using the following command

sudo blkid | grep crypto_LUKS
/dev/sda1: UUID="<crypt_device_uuid>" TYPE="crypto_LUKS" PARTUUID="*****"

Get device_uuid using the following command

sudo blkid | grep ext4
/dev/mapper/pd1: UUID="<device_uuid>" BLOCK_SIZE="4096" TYPE="ext4"

Validate the disk with lsblk to list block devices.

lsblk -f
NAME FSTYPE FSVER LABEL UUID FSAVAIL FSUSE% MOUNTPOINTS
sdb
└─sdb1 crypto_LUKS 2 `crypt_device_uuid xxxx-xxxx-xxxx-xxx-xxxx`
└─disk_name_mount_point ext4 1.0 `device_uuid like xxxx-xxxx-xxxx-xxx-xxxx` 1.9T 54% /mnt/disk_name_mount_point

Format the entry below.

{
"password": "luks password",
"crypt_device_uuid": "UUID of the device when type is `crypto_LUKS` like xxxx-xxxx-xxxx-xxx-xxxx",
"luks_key_base64": "Base 64 encoded LUKS key",
"device_uuid": "UUID of the device when type is `ext4` like xxxx-xxxx-xxxx-xxx-xxxx"
}

Crypt disk will be mounted using crypt_device_uuid in /etc/crypttab at /dev/mapper/{{ disk name mount point }}.

$ cat /etc/crypttab

{{ disk name mount point }} UUID={{ crypt_device_uuid }} /etc/cryptsetup-keys.d/{{ disk name mount point }} luks,nofail,timeout=300

And then /dev/mapper/{{ disk name mount point }} will be mounted at /mnt/{{ disk name mount point }} using /dev/mapper/{{ disk name mount point }} in /etc/fstab.

cat /etc/fstab

/dev/mapper/{{ disk name mount point }} /mnt/{{ disk name mount point }} ext4 rw,relatime,nofail,x-systemd.requires=/dev/mapper/{{ disk name mount point }},x-systemd.device-timeout=600 0 2

Before running patch_luks_ext4_mount, manually create a file at the root of the disk named .ansible_test_nfs_mount containing exactly Manually Created and Validated On a Working Disk. (trailing whitespace is trimmed). After mounting, the playbook reads that file back and fails if it's missing or its content doesn't match, confirming the correct disk was mounted.

Applications​

Application configurations

  • key: {{ cs_project_code }}/application-deployer/clusters/{{ Lab Name or Cluster name }}/hosts/{{ hostname }}/apps/{{ application }}/config

Application downstream dependency

  • key: {{ cs_project_code }}/application-deployer/clusters/{{ Lab Name or Cluster name }}/hosts/{{ hostname }}/apps/{{ application }}/generated/{{ resource }}

Variables​

OptionTypeDescriptionDefault
cs_vm_dockerd_tcp_socket_portintDocker TLS TCP socket portFrom Vault (hosts/{{ inventory_hostname }})
cs_glances_docker_imagestringGlances docker image{{ cs_vm_artifact_registry_containers_home }}/glances
cs_glances_docker_tagstringGlances docker tagubuntu-4.5.7-full
cs_glances_container_namestringGlances container nameglances
cs_glances_container_rootstringHost directory holding the Glances config and SSL certs/app/glances
cs_glances_web_api_portintGlances web server / REST API port61208
cs_glances_monitoring_passwordstringPassword protecting the Glances web server / REST APIFrom Vault (hosts/{{ inventory_hostname }})
cs_patch_nvidia_verify_docker_imagestringImage for the NVIDIA GPU passthrough check{{ cs_vm_artifact_registry_containers_home }}/ubuntu
cs_patch_nvidia_verify_docker_tagstringTag for the NVIDIA GPU passthrough check image26.04
cs_vm_ssh_service_user_idstringService user for SSHcloudinit
cs_restic_versionstringRestic version to install0.19.1
cs_vm_attached_luks_ext4_diskslistList of LUKS disks to mount[]
cs_nfs_server_mountlistList of NFS exports[]
cs_nfs_client_mountlistList of NFS imports[]
cs_nfs_server_portintNFS server (nfsd) port2049
cs_nfs_server_rpc_portintrpcbind / portmapper port111
cs_nfs_server_mountd_portintmountd port20048
cs_nfs_server_statd_portintstatd (NSM) port32765
cs_nfs_server_lockd_portintlockd (NLM) TCP and UDP port32803
cs_nfs_server_udpboolEnable NFS over UDPfalse
cs_nfs_server_tcpboolEnable NFS over TCPtrue
cs_nfs_server_vers3boolEnable NFSv3true
cs_nfs_server_vers4boolEnable NFSv4true
cs_nfs_server_vers4_0boolEnable NFSv4.0true
cs_nfs_server_vers4_1boolEnable NFSv4.1true
cs_nfs_server_vers4_2boolEnable NFSv4.2true
cs_vm_ssh_service_user_gidintService user GID1996
cs_vm_ssh_service_user_uidintService user UID1997
cs_vm_ssh_service_user_passwordstringService user passwordFrom Vault (hosts/{{ inventory_hostname }})
cs_vm_ssh_service_keyfilestringSSH key file path{{ playbook_dir }}/.ansible/.ssh/id_rsa_home_lab-{{ inventory_hostname }}
cs_vm_ssh_machine_known_hosts_filestringPer-host known_hosts file path, used as ansible_ssh_common_args's UserKnownHostsFile{{ playbook_dir }}/.ansible/.ssh/known_hosts-{{ inventory_hostname }}
cs_vm_ssh_machine_public_keyslistThe host's own SSH public keys, one known_hosts entry written per keyFrom Vault (hosts/{{ inventory_hostname }})
cs_vm_artifact_registry_schemastringArtifact registry URL schemeFrom DevOps Server
cs_vm_artifact_registry_netlocstringArtifact registry hostnameFrom DevOps Server
cs_vm_artifact_registry_containers_homestringContainer image namespace (host + project path)From DevOps Server
cs_vm_artifact_registry_generic_homestringGeneric artifact registry base URLFrom DevOps Server
cs_vm_artifact_registry_userstringArtifact registry userFrom DevOps Server
cs_vm_artifact_registry_passwordstringArtifact registry passwordFrom DevOps Server
cs_restic_binstringRestic binary path/usr/local/bin/restic
cs_restic_checksum_mapdictChecksum for the downloaded archive, keyed by ansible_facts.architecturePer version/arch map
cs_restic_download_urlstringArtifact registry URL for the restic binaryConstructed from version/arch
cs_restic_download_deststringLocal download destination for the restic archive/tmp/restic_{{ cs_restic_version }}_linux_{{ ansible_facts.architecture }}.bz2

cs_nfs_server_mount​

cs_nfs_server_mount:
- source_directory: /mnt/external-hdd-15
client_ip_cidr: '<cluster CIDR>' # the cluster CIDR read from Vault

cs_nfs_client_mount​

cs_nfs_client_mount:
- inventory_hostname: sash-m1
source_directory: /mnt/external-hdd-15
mount_path: /mnt/external-hdd-15
mount_options: defaults # optional, defaults to `defaults`; `port`, `mountport` and the systemd device timeout are appended

Each exported source_directory must contain a manually created file named .ansible_test_nfs_mount with the exact content Manually Created and Validated On a Working Disk.: the client mount task reads this file back through the mount and fails if it's missing or its content doesn't match, confirming the share mounted correctly.

Tasks​

System patching (patch_system)​

  • Installs basic packages (python3, procps, bash).
  • Installs the PostgreSQL client (postgresql-client, python3-psycopg2, libpq-dev) so every patched host can run psql/pg_dump/pg_restore and the community.postgresql Ansible modules without any service-specific task installing it itself.
  • Adds service user and configures SSH access.
  • Performs full Linux patching and sets hostname.
  • Hardens SSH and configures fail2ban.
  • Removes service user from the wheel group (if present).
  • Creates /app directory.

Restic installation (patch_install_restic)​

  • Removes any apt-installed restic and installs dependencies (rsync, fuse, bzip2, pigz).
  • Downloads the Restic binary from the internal generic artifact registry (cs_vm_artifact_registry_generic_home), authenticated with cs_vm_artifact_registry_user / cs_vm_artifact_registry_password, and installs it to /usr/local/bin/restic, then asserts the installed version matches cs_restic_version.

LUKS mount (patch_luks_ext4_mount)​

  • Installs cryptsetup.
  • Mounts LUKS encrypted ext4 disks specified in cs_vm_attached_luks_ext4_disks.
  • Reads .ansible_test_nfs_mount from the mount point and asserts its content is exactly Manually Created and Validated On a Working Disk., failing if the file is missing or the content doesn't match.

Docker configuration (patch_docker_socket)​

  • Installs Docker using geerlingguy.docker role.
  • Configures Docker to listen on a TLS-secured TCP socket.
  • Generates and signs certificates using the internal Root CA.
  • Serves the full certificate chain (leaf plus Root CA content) to TLS clients, so the socket verifies against a CA file holding only the root certificate.
  • Configures Docker login for the private registry.
  • Opens the Docker TLS TCP socket port in UFW.
  • Generates a short-lived client certificate and verifies the socket by calling its /version endpoint with mutual TLS.

NFS server (patch_nfs_server)​

  • Installs and configures NFS kernel server.
  • Exports directories specified in cs_nfs_server_mount.
  • Configures /etc/nfs.conf (the modern nfs-utils config):
    • Sets the nfsd port and the protocol/version toggles (udp, tcp, vers3, vers4, vers4.0, vers4.1, vers4.2).
    • Pins the mountd, statd (NSM) and lockd (NLM) ports so the NFSv3 RPC services can be firewalled.
    • Overrides the rpcbind.socket systemd unit to explicitly pin the port to cs_nfs_server_rpc_port (the standard 111) rather than leaving it to the unit default.
    • Removes the legacy RPCNFSDARGS / RPCMOUNTDOPTS lines from /etc/default/nfs-kernel-server.
  • Opens the nfsd, rpcbind, mountd, statd and lockd ports (cs_nfs_server_port, cs_nfs_server_rpc_port, cs_nfs_server_mountd_port, cs_nfs_server_statd_port, cs_nfs_server_lockd_port) in UFW, on both TCP and UDP.
  • Enables and starts rpcbind and nfs-server services.

NFS client (patch_nfs_client)​

  • Installs nfs-common.
  • Mounts NFS shares specified in cs_nfs_client_mount, passing both port and mountport mount options from the server's hostvars to support non-standard NFS ports.
  • Reads .ansible_test_nfs_mount from each mounted share and asserts its content is exactly Manually Created and Validated On a Working Disk., failing if the file is missing or the content doesn't match.

Nvidia (patch_nvidia)​

  • Installs linux-headers-<kernel> for the running kernel and pciutils (for lspci), then runs lspci and inspects its output for an NVIDIA card: the Ansible equivalent of lspci | grep -E "(VGA|3D)" | grep -E "(NVIDIA)".
  • If no NVIDIA GPU is detected, the install is skipped.
  • If a GPU is detected, the open-source nouveau driver is dealt with first: the xserver-xorg-video-nouveau package is purged, nouveau is blacklisted (/etc/modprobe.d/blacklist-nouveau.conf), the initramfs is rebuilt, and the play verifies nouveau is not currently loaded; if it still is, the play fails asking for a reboot before re-running.
  • Once nouveau is clear, NVIDIA's own CUDA repository is added so nvidia-driver comes from NVIDIA rather than the Debian repos: the cuda deb822 repository (signed-by the embedded NVIDIA CUDA signing key) is added, pointed at https://developer.download.nvidia.com/compute/cuda/repos/debian<major-version>/<arch>/. This deliberately avoids NVIDIA's cuda-keyring_*.deb installer package in favor of an auditable, idempotent Ansible-managed source. The apt cache is cleaned afterwards.
  • The nvidia-driver package is then installed (it bundles nvidia-smi).
  • A reboot is required after a driver install to load the kernel module; the play surfaces a notice but does not reboot automatically.
  • DRM KMS mode-setting is enabled (options nvidia-drm modeset=1 in /etc/modprobe.d/nvidia-drm-modeset.conf), rebuilding the initramfs when the file changes; without it, Wayland compositors (e.g. kwin_wayland under KDE Plasma/SDDM) can't find a usable DRM device and the session fails to start. A reboot is required for this to take effect; the play surfaces a notice but does not reboot automatically.
  • The host driver is verified with a native nvidia-smi. On the run that first installs the driver this is expected to fail (module not loaded until reboot) and the play stops before the container toolkit, asking for a reboot. On a later run the driver must already be loaded; a failure there aborts the play.
  • Once nvidia-smi confirms the driver is working, and before any Docker/toolkit work, the play hard-fails if either check below doesn't pass (nvidia-smi only proves the display/compute driver loaded, not that GPU passthrough into containers will work):
    • lsmod output must contain a loaded nvidia_uvm module.
    • Both /dev/nvidia-uvm and /dev/nvidia-uvm-tools device files must exist.
  • When a GPU is present, the driver is loaded, and Docker is installed (service_facts), the NVIDIA Container Toolkit is set up:
    • Downloads and dearmors the upstream signing key into /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg.
    • Adds the nvidia-container-toolkit deb822 repository (signed-by the keyring above).
    • Installs nvidia-container-toolkit, nvidia-container-toolkit-base, libnvidia-container-tools, and libnvidia-container1.
    • Runs nvidia-ctk runtime configure --runtime=docker to register the NVIDIA runtime in /etc/docker/daemon.json, and restarts Docker only when that config actually changes (compared by checksum).
    • Pulls the latest ubuntu:26.04 verification image from the private artifact registry (cs_patch_nvidia_verify_docker_image / cs_patch_nvidia_verify_docker_tag, via community.docker.docker_image_pull), which Docker is logged into by the patch_docker_socket stage. The play fails if the image can't be pulled.
    • Verifies GPU passthrough by running nvidia-smi in a throwaway ubuntu:26.04 container with --runtime=nvidia --gpus all (via community.docker.docker_container), asserting it exits cleanly.
    • If Docker is absent, the toolkit step is skipped.

Requires contrib non-free non-free-firmware components on the Debian repository (see prerequisites) for nvidia-driver.

Glances (patch_glances)​

  • Ensures Docker is running (installed by patch_docker_socket).
  • Detects whether an NVIDIA GPU is usable on the host by running nvidia-smi (non-fatal; hosts without a working driver skip this); if patch_nvidia has installed and loaded the driver, the GPU is attached to the container; otherwise Glances runs without it.
  • Creates cs_glances_container_root (/app/glances) and a certs subdirectory on the host.
  • Generates and signs an SSL server certificate using the internal Root CA, and writes a full-chain certificate (leaf plus Root CA content) alongside the private key.
  • Writes glances.conf in cs_glances_container_root (ssl_certfile / ssl_keyfile pointing at the mounted cert paths, passwords.local_password_path pointing at the mounted config directory so Glances picks up the pre-generated password file below instead of prompting, and fs.hide excluding /etc/resolv.conf, /etc/hosts, and /etc/hostname; with network_mode/pid_mode set to host, Glances' FS plugin picks these up as filesystem mount points from other containers on the host, which crashes the plugin and, with it, any REST API response bundled with FS in the same request, including the network/IP data used by the web UI).
  • Writes glances.pwd in cs_glances_container_root, containing cs_glances_monitoring_password hashed with the glances_password_hash filter, which replicates Glances' own PBKDF2-SHA256 password-file format. Glances' --password flag reads this file directly rather than prompting interactively (which would hang/crash the non-interactive container).
  • Opens cs_glances_web_api_port in UFW.
  • Starts the Glances container (network_mode: host, pid_mode: host) with the web server (-w) enabled on cs_glances_web_api_port and password protection (--password) turned on, so it enforces the password file written above, mounting the Docker socket (read-only), / at /rootfs (read-only), and cs_glances_container_root at /glances/conf (read-only). When an NVIDIA GPU was detected, the container is started with --runtime=nvidia --gpus all (via runtime / device_requests) and NVIDIA_VISIBLE_DEVICES / NVIDIA_DRIVER_CAPABILITIES set so the Glances GPU plugin can report it.
  • Verifies the container by calling its /api/4/status endpoint over HTTPS, validated against the Root CA.

systemd cleanup (dangerously_cleanup_systemd)​

Destructive. Only runs when this exact tag is passed explicitly (it is intentionally excluded from patch and every other patch_* tag) and tears down the host-level services this play installs, in order:

  • Removes the Glances container, closes its UFW port (cs_glances_web_api_port), and deletes cs_glances_container_root (/app/glances).
  • Deletes the Docker systemd override directory (/etc/systemd/system/docker.service.d) and reloads the systemd daemon. Docker itself stays installed.
  • Stops and disables nfs-client.target on hosts with a non-empty cs_nfs_client_mount (the mounts themselves and their /etc/fstab entries are left in place).
  • On hosts with a non-empty cs_nfs_server_mount, stops and disables nfs-server, rpcbind.socket, and rpcbind, deletes /etc/exports, reverts every patch_nfs_server-managed key in /etc/nfs.conf, deletes /etc/systemd/system/rpcbind.socket.d, reloads the systemd daemon, and closes the NFS/rpcbind UFW rules (cs_nfs_server_port, cs_nfs_server_rpc_port, cs_nfs_server_mountd_port, cs_nfs_server_statd_port, cs_nfs_server_lockd_port, and the named nfs UFW application rule). nfs-kernel-server/nfs-common are left installed.
  • For every disk in cs_vm_attached_luks_ext4_disks: unmounts /mnt/<disk> and removes it from /etc/fstab, stops and disables systemd-cryptsetup@<disk>.service, closes the LUKS mapping (cryptsetup close <disk>), removes the disk's line from /etc/crypttab, reloads the systemd daemon, and deletes the local keyfile at /etc/cryptsetup-keys.d/<disk>. The Vault-stored LUKS key material for the disk is never touched, so the disk stays decryptable elsewhere; the /mnt/<disk> mount point directory itself is left in place.

Set default target to multi-user (patch, patch_system, patch_install_restic, restic_backup, patch_luks_ext4_mount, patch_docker_socket, patch_nfs, patch_nfs_server, patch_nfs_client, patch_nvidia, patch_glances)​

  • Systemd: sets the default boot target to multi-user.target (systemctl set-default multi-user.target).
  • Runs whenever any patch tag is passed, not just patch_nvidia.

Reboot (patch, patch_system, patch_install_restic, restic_backup, patch_luks_ext4_mount, patch_docker_socket, patch_nfs, patch_nfs_server, patch_nfs_client, patch_nvidia, patch_glances)​

  • Reboots the patched host using the ansible.builtin.reboot module.
  • Detects whether the play is executing inside Gitea Actions or GitHub Actions by reading the GITEA_ACTIONS and GITHUB_ACTIONS environment variables on the controller with lookup('ansible.builtin.env', ...).
  • Only reboots when neither variable is true: running under CI skips the reboot, since the CI runner itself can be the patched host. It is the last task in the Patch play.

Tags​

  • patch: Run all patching tasks.
  • patch_system: Only run system patching.
  • patch_install_restic: Only install Restic.
  • patch_luks_ext4_mount: Only mount LUKS disks.
  • patch_docker_socket: Only configure Docker TLS socket.
  • patch_nfs: Configure both the NFS server and the NFS client.
  • patch_nfs_server: Only configure NFS server.
  • patch_nfs_client: Only configure NFS client.
  • patch_nvidia: Only detect and install the NVIDIA driver.
  • patch_glances: Only deploy the Glances monitoring container.
  • dangerously_cleanup_systemd: Destructive. Removes Glances, the Docker systemd override, and the NFS client/server and LUKS ext4 (cs_vm_attached_luks_ext4_disks) setup, including their systemd units, config, and (for LUKS) local keyfiles. Never included by patch or any other patch_* tag.

CI workflow​

The workflow (patch.yml) runs on metal runners (bare-metal hosts, not Docker). Jobs are structured as a single sequential chain, each depending on the one before it via needs:

  • patch-system: system patching (runs first)
  • patch-install-restic: restic installation (after patch-system)
  • patch-luks-ext4-mount: LUKS disk mounting (after patch-install-restic)
  • patch-nfs-server: NFS server (after patch-luks-ext4-mount)
  • patch-nfs-client: NFS client (after patch-nfs-server)
  • patch-docker-socket: Docker TLS socket (after patch-nfs-client)
  • patch-nvidia: NVIDIA driver install (after patch-docker-socket)
  • patch-glances: Glances monitoring container (after patch-nvidia)

Each job runs its own ansible-playbook invocation scoped to a single tag, and a job only runs once the job it needs has completed successfully.

Deployment​

uv sync --all-extras --all-packages --no-progress
uv --offline run --no-sync --no-progress ansible-galaxy install -r requirements.yml
uv --offline run --no-sync --no-progress ansible-playbook playbook.yml --tags patch