# Issues with fs\_lvm

**URL:** <https://forum.opennebula.io/t/issues-with-fs-lvm/3425>\
**Category:** Integration Support\
**Created:** [December 7, 2016, 7:17pm UTC](https://forum.opennebula.io/t/issues-with-fs-lvm/3425 "2016-12-07T19:17:40Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![olivier.toupin](https://yyz1.discourse-cdn.com/flex031/user_avatar/forum.opennebula.io/olivier.toupin/32/4968_2.png) [@olivier.toupin](https://forum.opennebula.io/u/olivier.toupin)\
**Post date:** [December 7, 2016, 7:17pm UTC](https://forum.opennebula.io/t/issues-with-fs-lvm/3425/1 "2016-12-07T19:17:40Z")

</div>

Hi,

I wanted to share to issues we experienced with fs\_lvm, and how we mitigated it.

### Problems

1. VM migration was broken. The datastore is expected to be shared, even if it the documentation say it doesn’t need to be. And if we share it, then volatile disks are stored on the NFS host, not on the local machine.

2. Upon host restart the VMs couldn’t reboot. The VMs directory structure was still there and in good shape, but the LV was no longer activated, causing the VM to fail to boot.

3. fs\_lvm force raw image conversion from qcow2 during clone. This would break thin provisioning (if our SAN supported it) and also causing longer “prolog” since, the full disk (filled with zero) would be copied instead of than the very small image (8GB vs 600MB, in our case)

4. Since the image store “TM\_MAD” is used to trigger the lvm “copy” we need two image stores. One standard for everything else (local storage, etc.) and a second one just for fs\_lvm. And this is even if the images are not really stored on LVM but are both on local storage anyway.

### Mitigation

1. Modified the MV script to actually copy (via rsync + ssh) the symlinks and volatile disks before activating the LV.

2. Added a hook on VM start, to double-check that the LV is activated, if a LV symlink is present and but broken.

3. Modified the clone script to use dd instead of qemu-img convert. Also, we run qemu-resize to fix the qcow image size.

4. We didn’t fix this one yet, but I included it anyway. Looking at the code, I understand why this is. It would seem natural, however, than if we select an SSH or Shared image and we select the VM to be deployed on a “fs\_lvm” datastore, it should work. Do you have any plan on improving this? We thought of making a new “wrapper” TM, say “ssh+fs\_lvm”, that would wrap SSH and FS\_LVM together and trigger the real TM according to the destination DS TM .

### Finally

- Here are the added and modified files ([gist](https://gist.github.com/oliviertoupin/69657df6c35c467669cffa77639e9cdc))

- Those changes will work for us on the short term. But any idea if those could get included upstream? Or you have other solutions for the future to fix these issues?

- Which of those should considered bugs, and filed in the bug tracker?

- Hopefully, those tricks can be used by other organizations running with the same problems.

---

<div class="post-metadata">

**Author:** ![jfontan](https://yyz1.discourse-cdn.com/flex031/user_avatar/forum.opennebula.io/jfontan/32/18_2.png) [@jfontan](https://forum.opennebula.io/u/jfontan)\
**Post date:** [December 15, 2016, 4:07pm UTC](https://forum.opennebula.io/t/issues-with-fs-lvm/3425/2 "2016-12-15T16:07:39Z")

</div>

> [@olivier.toupin](#):
>
> **Problems**
> 
> VM migration was broken. The datastore is expected to be shared, even if it the documentation say it doesn’t need to be. And if we share it, then volatile disks are stored on the NFS host, not on the local machine.

I have opened a ticket to change the documentation and state that system datastore needs NFS:

> **[Bug #4950: Change fs\_lvm documentation to warn about the need of shared...](https://dev.opennebula.org/issues/4950.html)**
>
> Redmine

We can take a looks about the volatile disks problem. Maybe we can alleviate the problem creating volatile disks. Would this be enough for your use case?

> [@olivier.toupin](#):
>
> Upon host restart the VMs couldn’t reboot. The VMs directory structure was still there and in good shape, but the LV was no longer activated, causing the VM to fail to boot.

You’re right and thanks for the hook. I’ve opened another issue for this:

> **[Bug #4951: Reactivate LV's on rebooted hosts - OpenNebula - OpenNebula...](https://dev.opennebula.org/issues/4951.html)**
>
> Redmine

This one is a bit tricky as we would like to have it integrated in the driver itself. We will check what could be the best way to implement activation for already created VMs.

> [@olivier.toupin](#):
>
> fs\_lvm force raw image conversion from qcow2 during clone. This would break thin provisioning (if our SAN supported it) and also causing longer “prolog” since, the full disk (filled with zero) would be copied instead of than the very small image (8GB vs 600MB, in our case)
> 
> Modified the clone script to use dd instead of qemu-img convert. Also, we run qemu-resize to fix the qcow image size.

We believe that most people use lvm for performance reasons so writing qcow2 format into it doesn’t make that much sense. I’ll be checking if thin LVs could be used and if they fix the problem.

Moreover, I think that resized qcow2 inside the LV can cause problems when the disk is almost full as qcow2 also stores metadata so the total file size for a full disk will be higher that LV size.

> [@olivier.toupin](#):
>
> We didn’t fix this one yet, but I included it anyway. Looking at the code, I understand why this is. It would seem natural, however, than if we select an SSH or Shared image and we select the VM to be deployed on a “fs\_lvm” datastore, it should work. Do you have any plan on improving this? We thought of making a new “wrapper” TM, say “ssh+fs\_lvm”, that would wrap SSH and FS\_LVM together and trigger the real TM according to the destination DS TM .

Can you please elaborate on this? We don’t fully understand what you are proposing.

> [@olivier.toupin](#):
>
> Here are the added and modified files

Thanks! I think this could be the start of a document on how to configure fs\_lvm for non shared system datastores.

> [@olivier.toupin](#):
>
> Those changes will work for us on the short term. But any idea if those could get included upstream? Or you have other solutions for the future to fix these issues?

We are going to check if there is a way to support both shared and non shared configurations of system datastore.
