Completed missionOn-premises virtualisation with Proxmox

On-premises virtualisation with Proxmox for a company in the medical sector
Sector
BioTech
Engagement duration
1 week

Mission summary

Cloud Algebra helped a company in the medical sector add a virtualisation layer (Proxmox) on top of its physical infrastructure, improving utilisation rates, sizing, access control, isolation, backups, business continuity and disaster recovery…

Engagement details

Context

The client we supported runs its own physical infrastructure with:

  • a local LAN unreachable from the outside;
  • several matrix-computing nodes used to train AI models;
  • other physical servers storing data;
  • applications and infrastructure building blocks.

Challenges

Before our engagement, all of the client’s data and applications were installed directly on the physical machines. This raises a few challenges:

  • How to use a server’s full hardware capacity while avoiding a “catch-all” effect, i.e. a server hosting data, applications and services that are too heterogeneous, for lack of a better home?
  • How to properly isolate the data and applications hosted on the same server so that the vulnerabilities (security, resilience and performance) of one do not spread to the others?
  • How to guard against hardware failures and avoid prolonged outages?
  • How to guard against the risk of data loss?

Introducing Proxmox

Proxmox is an open-source type-1 hypervisor based on Debian (Linux) that makes it easy to virtualise a physical server into several isolated virtual servers.

With Proxmox, several virtual machines (VMs) and LXC containers can run on the same physical infrastructure, whether a single machine or a cluster of several. Physical and virtual machines in the cluster are managed through an ergonomic web interface or directly with command-line tools.

Proxmox natively includes many key features:

  • Automated or manual backups and restores;
  • Live VM snapshots;
  • Replication and failover between nodes, enabling high availability;
  • Fine-grained resource allocation (CPU, RAM, storage);
  • Integration with various storage back-ends (PBS, LVM, ZFS, Ceph, etc.).

Why Proxmox

After auditing current and projected usage, Cloud Algebra recommended introducing Proxmox on a pilot machine and then, subject to satisfaction, rolling it out gradually across the whole infrastructure as a cluster of machines.

Virtualising the client’s physical resources into virtual machines and logical volumes is a first strategic step towards improving the company’s resilience and security. A well-configured Proxmox installation effectively addresses the challenges listed above:

  • Proxmox enables better isolation and better sizing of applications;
  • Proxmox secures sensitive operations (updates, migrations or tests) through very convenient backups, snapshots and restores;
  • Proxmox harmonises the management of the various production systems;
  • Proxmox simplifies capacity and workload management thanks to clustering, horizontal node scaling, VM migration and vertical VM scaling;
  • Proxmox greatly simplifies VM backups and restores for a fast response to incidents.

Understanding that installing Proxmox anticipated and prevented many risks the client was not yet covered against, the client gladly accepted the pilot project to get to know the solution and possibly deploy it more widely.

Cloud Algebra also recommended drawing up a Business Continuity Plan and a Disaster Recovery Plan in which Proxmox can be leveraged to restore VMs on another host, or the infrastructure on another site.

Engagement: the pilot project

Summary

Cloud Algebra migrated the data and applications hosted on an old physical server to a new machine on which Proxmox was installed and configured. Proxmox was set up so that applications that used to share a single machine are now isolated in their own dedicated VM and kept under control, bringing more order, flexibility, resilience and security to the client.

VMs

We created five VMs (applications anonymised):

  • vm-app-1 (Ubuntu Server 24.04 LTS, 6 vCPU, 4 GB RAM, static IP)
  • vm-app-2 (Ubuntu Server 24.04 LTS, 2 vCPU, 2 GB RAM, static IP)
  • vm-app-3 (Ubuntu Server 24.04 LTS, 1 vCPU, 2 GB RAM, static IP)
  • vm-sandbox (Ubuntu Server 24.04 LTS, 1 vCPU, 2 GB RAM, static IP)
  • vm-pbs (Proxmox Backup Server 4.0, 2 vCPU, 4 GB RAM, static IP)

The VMs are now isolated, and a crash or corruption of one application has no impact on the applications in the other VMs. The capacity allocated to each VM can also be changed easily, subject to available capacity, so that sizing follows each application’s workload.

Each VM hosts at most one application, its dependencies and its data, nothing more. Any new application will be deployed in its own VM so that this rule is respected and the infrastructure stays maintainable. The vm-sandbox is a non-persistent VM that lets the technical teams experiment and run proofs of concept without any risk to production applications.

For each VM, an appropriate backup strategy was put in place, taking into account the criticality of the application, its RTO (Recovery Time Objective) and its RPO (Recovery Point Objective). The vm-pbs hosts the backup server.

Storage

The pilot server has 2 physical disks:

  • 1 SSD of 400+ GB
  • 1 HDD of 15+ TB

Both disks are registered as LVM Physical Volumes in the Proxmox node:

  • /dev/sda for the SSD
  • /dev/sdb for the HDD

Each Physical Volume was added to the “pve” LVM Volume Group dedicated to the Proxmox Virtual Environment.

Logical Volumes were also added:

  • Proxmox’s default Logical Volumes — data, root, swap — all three on the /dev/sda Physical Volume (SSD);
  • Three specific Logical Volumes on the HDD:
    • storage-app-1: 2 TB, meant to hold all the external data of a self-hosted application on a dedicated volume; attached to vm-app-1 and mounted on /mnt/storage-app-1 via /etc/fstab;
    • storage-assets: 8 TB, meant to hold any other data clearly distinct from the application (the company’s raw data); attached to vm-sandbox and automatically mounted on /mnt/storage-assets via /etc/fstab;
    • storage-backup: 2 TB, meant to hold Proxmox backups generously; attached to vm-pbs and automatically mounted on /mnt/storage-backup via /etc/fstab.

Note that adding two more HDDs would allow replacing the LVM storage layout with a ZFS RAIDZ-5 layout on the HDDs with an Intent Log and a cache on the SSD. These changes would considerably reduce the risk of data loss and allow instant VM snapshots thanks to Copy-on-Write, while making better use of the SSD. This is the configuration that will be implemented for the second Proxmox node.

Backups with PBS: Proxmox Backup Server

Cloud Algebra set up an automated backup infrastructure with “PBS”. Backups are daily, compressed, encrypted and deduplicated to maximise the security of sensitive data sent off-site and minimise external storage needs.

On PBS, user permissions were restricted accordingly by creating a “backup-operator” user with only “DatastoreAdmin” rights on the relevant datastores and “RemoteAdmin” rights on the relevant remotes.

Two copies are kept: a local copy, itself replicated to external S3 storage that Cloud Algebra ordered and configured with the client to assist them in the process. To set up these two copies, a first local datastore was created in a fairly standard way and configured as the target datastore in PVE through the PVE x PBS integration. Then an S3 endpoint and a second datastore were created on the PBS instance. Finally, replication was set up between the two datastores. The two datastores have different retention policies: only backups less than one month old (the most likely to be fully restored) are kept locally, while off-site backups up to 3 years old are kept — less likely to be restored, but no less important.

Besides backups, backup integrity verification jobs, pruning jobs and garbage-collection jobs were configured.

Roadmap: redundancy, HA, rebalancing and IaC

Rolling Proxmox out to all of the client’s servers would be a huge step up in infrastructure management maturity. Cloud Algebra plans to continue with the following work:

  • Installing a second Proxmox node and forming a cluster;
  • Disk redundancy on each node, in RAIDZ-5 with ZFS;
  • Migrating VMs to shared storage (Ceph or NFS) to enable live migration from one node to another;
  • Configuring High Availability to allow automatic failover when a node becomes unavailable;
  • Configuring a placement policy (affinity, anti-affinity, balance) to keep the load balanced at all times;
  • Scripting and reproducibility of the entire configuration to allow very fast redeployment in the event of a disaster.

Conclusion

By putting in place a solid, modern virtualisation foundation such as Proxmox, the client’s infrastructure becomes simpler to manage, more reliable and far more adaptable to changing needs. Resources are used better, backups are centralised and encrypted, and maintenance is greatly simplified.

This approach gives the client a solid, scalable and secure base, perfectly aligned with its performance and continuity requirements. A pragmatic, sustainable, forward-looking transformation.

All our projects

From plan to action

Ready to take the next step?Cloud Algebra secures, stabilises and optimises your cloud infrastructure.

Draw on our teams’ expertise to modernise, secure and optimise your information systems. Let’s talk about your project.

Work with us