How Should TM_MAD=local Be Used with Host-Local Datastores?

Hi!

I’m trying to deploy a super simple setup using OpenNebula 7.2:

  • one controller which hosts the base images
  • one cluster (“testcluster”) with one satellite (for now) with local storage that hosts the vms and keeps their corresponding disk images in /var/lib/one/datastores/ (which is a dedicated xfs-formatted mount point with enough capacity)
  • have a cache to avoid copying the base image to the satellite on every deployment
  • later, after adding additional hosts: being able to live migrate between the satellites

sounds pretty easy, doesn’t it? But somehow I’m stuck at using TM_MAD=local and/or my understanging is broken.

My satellite host is of type kvm, member of cluster “testcluster” (100), no special configuration.

I started with the following ds:

TYPE = SYSTEM_DS
TM_MAD = local
DS_MAD = fs
DISK_TYPE = FILE
RESTRICTED_DIRS = /
SAFE_DIRS = /var/tmp /var/lib/one/datastores
CACHE_ENABLE = YES
CACHE_PATH = /var/lib/one/datastores/cache
DATASTORE_CAPACITY_CHECK = YES
STAGING_DIR = /var/lib/one/datastores/staging

in this example, the ds id is 107 and it doesn’t matter whether or not I enable the cache or redefine the staging dir or cache path.

Now I have some dead giveaways that this failed:

  1. on the satellite there’s no /var/lib/one/datastores/107 (it’s not a permission problem, see below)
  2. I cannot see the free disk capacity anywhere. I understand with TM_MAD=local it’s not shown in the datastore itself, but I also cannot see it with “onehost show” nor in the Hosts-tab in fireedge.

When using TM_MAD=ssh and adding my satellite to the BRIDGE_LIST I can see the correct disk capacity and the ds directory (/var/lib/one/datastores/$id) is created one the satellite.

What am I doing wrong or what am I missing?

$ onehost show 1
HOST 1 INFORMATION                                                              
ID                    : 1                   
NAME                  : mysatellite
CLUSTER               : testcluster              
STATE                 : MONITORED           
IM_MAD                : kvm                 
VM_MAD                : kvm                 
LAST MONITORING TIME  : 06/15 13:29:09      

HOST SHARES                                                                     
RUNNING VMS           : 0                   
MEMORY                                                                          
  TOTAL               : 124.7G              
  TOTAL +/- RESERVED  : 124.7G              
  USED (REAL)         : 18.6G               
  USED (ALLOCATED)    : 0K                  
CPU                                                                             
  TOTAL               : 3200                
  TOTAL +/- RESERVED  : 3200                
  USED (REAL)         : 0                   
  USED (ALLOCATED)    : 0                   

MONITORING INFORMATION                                                          
ARCH="x86_64"
CGROUPS_VERSION="2"
CPUSPEED="0"
HOSTNAME="mysatellite"
HYPERVISOR="kvm"
IM_MAD="kvm"
KVM_CPU_FEATURES="3dnowprefetch,abm,acpi,adx,aes,apic,arat,arch-capabilities,avx,avx2,avx512-vpopcntdq,avx512bitalg,avx512bw,avx512cd,avx512dq,avx512f,avx512ifma,avx512vbmi,avx512vbmi2,avx512vl,avx512vnni,bmi1,bmi2,clflush,clflushopt,clwb,cmov,cx16,cx8,dca,de,ds,ds_cpl,dtes64,erms,est,f16c,fb-clear,flush-l1d,fma,fpu,fsgsbase,fsrm,fxsr,gds-no,gfni,hle,ibrs-all,intel-pt,invpcid,invtsc,la57,lahf_lm,lm,mca,mce,md-clear,mds-no,mmx,monitor,movbe,mpx,msr,mtrr,nx,pae,pat,pbe,pcid,pclmuldq,pconfig,pdcm,pdpe1gb,pge,pku,pni,popcnt,pschange-mc-no,psdp-no,pse,pse36,rdctl-no,rdpid,rdrand,rdseed,rdtscp,rfds-no,rtm,sbdr-ssdp-no,sep,sgx,sgxlc,sha-ni,skip-l1dfl-vmentry,smap,smep,smx,spec-ctrl,ss,ssbd,sse,sse2,sse4.1,sse4.2,ssse3,stibp,syscall,tm,tm2,tsc,tsc-deadline,tsc_adjust,tsx-ctrl,umip,vaes,vme,vmx,vpclmulqdq,wbnoinvd,x2apic,xgetbv1,xsave,xsavec,xsaveopt,xsaves,xtpr"
KVM_CPU_MODEL="Icelake-Server-v2"
KVM_CPU_MODELS="486 486-v1 Broadwell-noTSX Broadwell-noTSX-IBRS Broadwell-v2 Broadwell-v4 Cascadelake-Server-noTSX Cascadelake-Server-v3 Cascadelake-Server-v4 Cascadelake-Server-v5 Conroe Conroe-v1 Denverton-v2 Denverton-v3 Haswell-noTSX Haswell-noTSX-IBRS Haswell-v2 Haswell-v4 Icelake-Server-noTSX Icelake-Server-v2 IvyBridge IvyBridge-IBRS IvyBridge-v1 IvyBridge-v2 Nehalem Nehalem-IBRS Nehalem-v1 Nehalem-v2 Opteron_G1 Opteron_G1-v1 Opteron_G2 Opteron_G2-v1 Penryn Penryn-v1 SandyBridge SandyBridge-IBRS SandyBridge-v1 SandyBridge-v2 Skylake-Client-noTSX-IBRS Skylake-Client-v3 Skylake-Client-v4 Skylake-Server-noTSX-IBRS Skylake-Server-v3 Skylake-Server-v4 Skylake-Server-v5 Westmere Westmere-IBRS Westmere-v1 Westmere-v2 core2duo core2duo-v1 coreduo coreduo-v1 kvm32 kvm32-v1 kvm64 kvm64-v1 n270 n270-v1 pentium pentium-v1 pentium2 pentium2-v1 pentium3 pentium3-v1 qemu32 qemu32-v1 qemu64 qemu64-v1"
KVM_MACHINES="pc-i440fx-rhel7.6.0 pc pc-q35-rhel9.6.0 q35 pc-q35-rhel8.6.0 pc-q35-rhel9.4.0 pc-q35-rhel8.5.0 pc-q35-rhel8.3.0 pc-q35-rhel7.6.0 pc-q35-rhel8.4.0 pc-q35-rhel9.2.0 pc-q35-rhel8.2.0 pc-q35-rhel9.0.0 pc-q35-rhel8.0.0 pc-q35-rhel8.1.0"
MEMORY_ENCRYPTION="NONE"
MODELNAME="Intel(R) Xeon(R) Gold 5315Y CPU @ 3.20GHz"
RESERVED_CPU=""
RESERVED_MEM=""
TOTALCPU="3200"
TOTALMEMORY="130768388"
VERSION="7.2.0"
VM_MAD="kvm"

NUMA NODES

  ID CORES                    USED FREE
   0 -- -- -- -- -- -- -- --  0    16
   1 -- -- -- -- -- -- -- --  0    16

NUMA MEMORY

 NODE_ID TOTAL    USED_REAL            USED_ALLOCATED       FREE    
       0 61.8G    13.4G                0K                   48.4G
       1 62.9G    6.7G                 0K                   56.3G

NUMA HUGEPAGES

 NODE_ID SIZE     TOTAL    FREE     USED    
       0 2M       0        0        0
       0 1024M    0        0        0
       1 2M       0        0        0
       1 1024M    0        0        0

WILD VIRTUAL MACHINES

NAME                                                      DEPLOY_ID  CPU     MEMORY

VIRTUAL MACHINES

  ID USER     GROUP    NAME                                                                                                                                                                          STAT  CPU     MEM HOST
$ onedatastore show 107
DATASTORE 107 INFORMATION                                                       
ID             : 107                 
NAME           : fs_system_ds 
USER           : oneadmin            
GROUP          : oneadmin            
CLUSTERS       : 100                 
TYPE           : SYSTEM              
DS_MAD         : -                   
TM_MAD         : local               
BASE PATH      : /var/lib/one//datastores/107
DISK_TYPE      : FILE                
STATE          : READY               

DATASTORE CAPACITY                                                              
TOTAL:         : -                   
FREE:          : -                   
USED:          : -                   
LIMIT:         : -                   

PERMISSIONS                                                                     
OWNER          : um-                 
GROUP          : u--                 
OTHER          : ---                 

DATASTORE TEMPLATE                                                              
ALLOW_ORPHANS="FORMAT"
CACHE_ENABLE="YES"
CACHE_PATH="/var/lib/one/datastores/cache"
DATASTORE_CAPACITY_CHECK="YES"
DS_LIVE_MIGRATE="YES"
DS_MIGRATE="YES"
DS_MIGRATE_SNAP="YES"
NAME="fs_system_ds"
PERSISTENT_SNAPSHOTS="YES"
RESTRICTED_DIRS="/"
SAFE_DIRS="/var/tmp /var/lib/one/datastores"
SHARED="NO"
STAGING_DIR="/var/lib/one/datastores/staging"
TM_MAD="local"
TYPE="SYSTEM_DS"

IMAGES         

Hi Philippe,

it looks normal and you are just seem to mix use cases.

For the sake of uniformity, let’s use a term KVM host (hypervisor host) for cluster members. That’s how they are refered to in the OpenNebula documentation. It took me a while to realize you are not talking about RedHat Satellite.

Let me quote your questions:

on the satellite there’s no /var/lib/one/datastores/107 (it’s not a permission problem, see below)

Datastore is initially non-existent on the KVM host. It’s created when needed and stays there afterwards. Hence, once you define even a minimal datastore config with template:

TM_MAD = local
TYPE = SYSTEM_DS

It gets initialized on the Frontend and will be populated to the KVM host on the first use. Just make sure the KVM host has enough capacity in the default location for creating new datastores.

I cannot see the free disk capacity anywhere. I understand with TM_MAD=local it’s not shown in the datastore itself, but I also cannot see it with “onehost show” nor in the Hosts-tab in fireedge.

That’s normal, because by default the Frontend monitors datastores only “here” from its perspective. This “here” might be a real local resource or a mounted remote (NFS) filesystem, it doesn’t matter as the Frontend just looks literally on its own directory tree. The only way to tell the Frontend that a resource is distant is to fill the BRIDGE_LIST, but beware. It’s primarily intended to monitor shared resources that Frontend has no direct access to (i.e. not in its own directory tree/filesystem). Let’s consider two examples of a shared datastore use:

  1. mapped (shared) to all hosts AND frontend - then the frontend needs no extra config and will show space stats “here” from its perspective. It sees mounted NFS share as “own” directory and reads stats properly.

  2. mapped (shared) to all hosts but NOT frontend - then the frontend by default will show wrong storage space stats, still showing “here” from its perspective so a local (most probably empty) directory. In order to tell it to show a shared resource stats, you need to fill BRIDGE_LIST which makes it to jump over SSH to the defined bridge host and look there “locally” in defined location, if that makes sense.

Setting BRIDGE_LIST on datastore by adding KVM hosts, each with local datastore, will result in an unpredictable behavior, technically bringing unrelated stats from different hosts and randomize the outcome… The reason why it works for you now is that you have only one KVM host and “there” (BRIDGE_LIST) is one and only. It will break once you add another host.

If you need to monitor individual hosts’ local storage, Prometheus+Grafana is a way to go.

Please note that TM_MAD=ssh driver is not recommended and left only for backwards compatibility. It’s an old version of “local” and should not be used in the new deployments. MOre info here

Also, if you plan to enable live migration, you need a shared storage anyway so maybe better to build it early in the project.

ok, found that I needed to add a file “.monitor” containing “local” or “ssh” in the corresponding datastore on the satellite.

Hi Damian,

thanks for the detailed answer!

You mentioned that live migration requires a shared storage. I saw that ONe uses “–copy-storage-all” in vmm/kvm, so CMIIW but local storage should not be a problem for live migration, right?

Hi Philippe, of course. It’s a dig-up topic but I’ve noticed your post only now.
OpenNebula uses a combination of tar and qemu’s blockcopy funciton to send all VM’s belongings to the target host. However, due to the nature of the process, it takes longer than in the shared environment and may be impacted by events not related to OpenNebula. Especially with big Vms with multi-gigabyte disks. Therefore, for production a shared datastore is usually the way to go. Even if we accept the slower process, there might be a “Convergence” Problem. If the VM is running a highly active database, it might write new data to the local disk faster than your network can transmit it to the destination. If the migration cannot converge, KVM will loop indefinitely unless you have MIGRATE_AUTO_CONVERGE enabled in your VM templates to throttle the VM’s CPU.