Introduction
Labtainers, the Docker-based cybersecurity lab framework developed by the Naval Postgraduate School, gives students isolated, repeatable hands-on environments. But by default, Labtainers is designed to run on a local machine, typically inside a VirtualBox or VMware VM. Getting it to run cleanly on an OpenStack cloud — especially through Exosphere, a user-friendly OpenStack client — surfaces a set of non-trivial compatibility problems.
This post walks through the core conflict I ran into while building a fully OpenStack-compatible Labtainer image accessible through Exosphere, and how I resolved it. The work was carried out and validated on Jetstream2, an NSF-funded research cloud.
Note: Jetstream2 isn't a stand-in or simulation of OpenStack — its operational software environment is built directly on OpenStack, and it's accessible via the standard OpenStack CLI and SDK. It does add an institutional layer on top (ACCESS/CILogon authentication, multi-region federation, a curated image library), but the underlying infrastructure is genuine OpenStack. That means an image built and validated for Jetstream2 is, by definition, an image built for OpenStack — this work was tested in a real, multi-tenant, production-grade OpenStack cloud, not an isolated lab setup.
The Core Problem: Two Docker Installs, One System
Labtainer ships with its own Docker installation — the lab containers are pre-configured to run against that specific Docker daemon.
Exosphere, on the other hand, uses cloud-init/Ansible-based “exouser” provisioning tasks to automatically install Docker on every instance it creates, then deploys its Guacamole stack (guacd, guacamole, and related proxy services) on top of it, enabling one-click, browser-based shell and desktop access.
When both systems land on the same image:
- Exosphere's Docker installation step collides with the Docker already installed by Labtainer, and fails.
- Even if that collision is forced through, Labtainer labs don't run correctly on the Docker environment Exosphere sets up, due to version and configuration mismatches.
It's a classic case of two independent provisioning mechanisms trying to manage the same resource.
The Fix: Integration, Not Handoff
The logic behind the fix is simple, though the implementation required care: completely disable Exosphere's own Docker installation, and instead integrate Exosphere's Guacamole containers into the Docker daemon that Labtainer already provides.
Concretely:
- First, I identified the specific commit that Exosphere's exouser tasks were running from, then forked those tasks to my own GitHub account.
- During provisioning, I pointed the instance to pull the exouser tasks from my own GitHub repository — customized specifically for Labtainer — instead of Exosphere's default source.
- Removed the Docker installation step entirely — Labtainer's Docker is already present, and installing a second one only causes conflicts.
- Reconfigured Exosphere's Guacamole container deployment to target Labtainer's existing Docker daemon instead of a freshly installed one.
The result: a single Docker daemon runs both the native Labtainer lab containers and Exosphere's Guacamole stack side by side, with no conflicts.
Step by Step: Building a 100% OpenStack-Compatible Image
Once the Docker/Guacamole integration was solved, the image still needed to behave like a proper cloud image on OpenStack. That didn't come down to a single step — it required tracking down and fixing a series of practical issues one by one. None of these were documented “known issues” anywhere — I found each one by repeatedly deploying the image on a real OpenStack cloud through Exosphere, reading the logs, and isolating the step that was failing. Here's that process, step by step.
(The Docker/containerd conflict was already covered above in “The Fix: Integration, Not Handoff” — resolved by disabling Exosphere's own Docker installation step.)
Step 1: Base OpenStack Packages
Packages needed so the image can talk to OpenStack's metadata service, handle SSH key injection, and resize disks on boot:
sudo apt update
sudo apt install -y \
cloud-init \
cloud-guest-utils \
qemu-guest-agent \
openssh-server
sudo systemctl enable qemu-guest-agent
sudo systemctl start qemu-guest-agent
sudo systemctl enable ssh
Step 2: Removing unattended-upgrades
Issue: unattended-upgrades was holding the apt lock in the background, which blocked apt operations during cloud-init/provisioning. In the logs, this showed up as provisioning scripts silently hanging on a “Could not get lock /var/lib/dpkg/lock-frontend” error.
Fix:
sudo systemctl stop unattended-upgrades
sudo systemctl disable unattended-upgrades
sudo apt remove -y unattended-upgrades
Step 3: Installing TurboVNC + VirtualGL
Issue: The vncpasswd command wasn't found, and there was a package conflict between the VNC repository files — which caused apt update failures in later stages of the Ansible-based provisioning.
Fix:
# Add repos
curl -fsSL https://packagecloud.io/dcommander/turbovnc/gpgkey | sudo tee /etc/apt/trusted.gpg.d/turbovnc.asc
echo "deb https://packagecloud.io/dcommander/turbovnc/any/ any main" | sudo tee /etc/apt/sources.list.d/turbovnc.list
curl -fsSL https://packagecloud.io/dcommander/virtualgl/gpgkey | sudo tee /etc/apt/trusted.gpg.d/virtualgl.asc
echo "deb https://packagecloud.io/dcommander/virtualgl/any/ any main" | sudo tee /etc/apt/sources.list.d/virtualgl.list
# Install
sudo apt update
sudo apt install -y turbovnc virtualgl dbus-x11
# Remove repo files (to avoid conflicts with Ansible)
sudo rm -f /etc/apt/sources.list.d/turbovnc.list
sudo rm -f /etc/apt/sources.list.d/virtualgl.list
sudo rm -f /etc/apt/trusted.gpg.d/turbovnc.asc
sudo rm -f /etc/apt/trusted.gpg.d/virtualgl.asc
# vncpasswd symlink
sudo ln -s /opt/TurboVNC/bin/vncpasswd /usr/local/bin/vncpasswd
The key detail here: remove the repo files after installing the packages, not before. If the repo files stay on the image, later provisioning steps (Exosphere's own apt update calls) run into conflicts.
Step 4: Enabling Serial Console
Issue: Instances launched through Exosphere were timing out during boot. Digging into the logs showed that Exosphere reads the instance's serial console output to track the boot process, but the image didn't have that output enabled.
Fix: Add serial console parameters to GRUB:
sudo sed -i 's/GRUB_CMDLINE_LINUX_DEFAULT=".*"/GRUB_CMDLINE_LINUX_DEFAULT="console=tty1 console=ttyS0,115200"/' /etc/default/grub
sudo sed -i 's/GRUB_CMDLINE_LINUX=".*"/GRUB_CMDLINE_LINUX="console=tty1 console=ttyS0,115200"/' /etc/default/grub
sudo update-grub
Step 5: Cleanup and Generalization (MUST BE THE LAST STEP)
Issue: Cloud-init was caching stale state from a previous boot (old instance-id, old machine-id, old SSH host keys), which meant it skipped first-boot tasks (SSH key injection, hostname configuration, etc.) on new instances. This also risked every instance derived from the image sharing the same SSH host keys and machine-id.
Fix: Right before turning the image into a reusable template — and only as the very last step:
# Clean cloud-init
sudo cloud-init clean --logs
sudo rm -rf /var/lib/cloud/*
# Reset machine ID
sudo truncate -s 0 /etc/machine-id
sudo rm -f /var/lib/dbus/machine-id
# Remove SSH host keys
sudo rm -f /etc/ssh/ssh_host_*
# Update initramfs
sudo update-initramfs -u -k all
# (Optional) Clear logs and history
sudo rm -f /var/log/cloud-init*.log
history -c
# Shut down
sudo shutdown -h now
Doing this step last is critical: no further changes should be made to the image after this point — the Glance image should be created directly from this powered-off disk. Otherwise the cleaned-up state gets recreated, and the image keeps carrying the same identity (machine-id, SSH host keys) into every new instance.
Why This Combination Hasn't Been Documented Before
In researching this, I found two related but distinct prior efforts:
- Olivier Berger's 2018 blog post (“Demo of displaying labtainers labs in a Web browser through Guacamole”) shows Labtainer labs displayed in a browser via Guacamole — but it's a standalone Docker + Guacamole setup, with no connection to Exosphere or OpenStack.
- Exosphere's official Guacamole support (GitLab MR !335) is built around installing a fresh Docker + Guacamole stack on every instance — it never considers integrating into a pre-existing third-party Docker daemon like Labtainer's.
So the specific approach here — disabling Exosphere's own Docker provisioning and integrating its Guacamole containers into a third-party Docker daemon instead — doesn't appear to be documented anywhere publicly, as far as I could find.
Why Validating on Jetstream2 Matters
I ran this work on Jetstream2 rather than an isolated test OpenStack deployment — a real production environment. That matters because Jetstream2's infrastructure is built directly on OpenStack, so an image that works there is, by definition, OpenStack-compatible. This isn't a “Jetstream2-specific fix” — an image that works cleanly even underneath Jetstream2's institutional layers (authentication, network policy, multi-region federation) will work just as well on a simpler, single-region OpenStack deployment. Having validated it in a real, multi-tenant production environment gives the solution an extra degree of confidence.
Conclusion
This work makes it possible to run Labtainer cybersecurity labs on an OpenStack cloud, accessed through Exosphere, with zero manual setup on the student's side — just one click, in a browser. It offers a scalable deployment model for teaching institutions and removes the requirement for students to install VirtualBox or VMware locally.
Next steps: publishing this work on GitHub and contributing it back to both the Exosphere and Labtainers projects.