Repository navigation
🐞 Bug: Newly added disk/net device not added to boot_devices #289
Description
Activity
Re-validated against HyperCore 9.7.7.226383 (4-node cluster), collection
main@34308d2,
ansible-core 2.16.19. Answering the issue's own "Check if this is still problem - it might be
already fixed": still reproduces. Keeping open, with a sharper description and the mechanism
below.Reproduction
Ran the exact three-task shape from the description — create a VM with 0 disks and 0 NICs, then add
1 disk + 1 NIC with both listed inboot_devices:TASK 1 create VM, 0 disks / 0 NICs -> disks=0 nics=0 boot_devices=0 TASK 2 add 1 disk + 1 NIC, both bootable -> boot_devices count = 1 (expected 2) -> the single entry is the NIC: {type: virtio, mac: 7C:4C:58:ED:D5:6F, vlan: 0, connected: true, ...} -> disks reports ['virtio_disk'] — the disk IS attached, just absent from boot orderSo the original report is exactly right: only the NIC lands in boot order.
It self-corrects on a second pass — which makes it easy to miss
Running the identical task again converges:
PASS 1 changed=True boot_devices=['virtio'] PASS 2 changed=True boot_devices=['virtio_disk', 'virtio'] PASS 3 changed=False boot_devices=['virtio_disk', 'virtio']This is the same two-pass shape as #30's original report. It still matters, because anyone who
runs their playbook once gets a VM whose disk is silently not bootable — no error, no warning,
andchanged=Truereports success. For a VM that is about to be booted from that disk, the failure
surfaces later and far from its cause.Workaround:
vm_boot_deviceswithstate: setfixes it immediately afterwards —
boot_devicesbecomes['virtio_disk', 'virtio']. So the boot-order write path itself is fine;
the problem is what thevmmodule feeds into it.Mechanism (code-verified on
main@34308d2)Three things compose:
1. The VM object used for boot order is captured too early.
ensure_present
(plugins/modules/vm.py:448-460) readsvm_before, then runs_set_disksand_set_nics, and
only then calls_set_boot_order(module, rest_client, vm_before, existing_boot_order)— passing the
object it captured before the devices were created.2. Unresolvable boot devices are dropped silently.
set_boot_devices_order
(plugins/module_utils/vm.py:769-776):def set_boot_devices_order(self, boot_items): boot_order = [] for desired_boot_device in boot_items: vm_device = self.get_vm_device(desired_boot_device) if not vm_device: continue # <-- silent drop, no error or warning boot_order.append(vm_device["uuid"]) return boot_order
3. This explains why the NIC survives and the disk does not. The NIC path does refresh the
object it is about to hand on (plugins/module_utils/vm.py:1606-1610):updated_virtual_machine_TEMP = VM.get_by_old_or_new_name(module.params, rest_client=rest_client) updated_virtual_machine = vm_before updated_virtual_machine.nics = updated_virtual_machine_TEMP.nics del updated_virtual_machine_TEMP # TODO are nics from vm_before used anywhere?
The disk path does not. When
ManageVMDisks.ensure_present_or_setis called from thevmmodule it
takes thereturn changedpath atvm.py:1450; theget_vm_by_namerefresh atvm.py:1443only
runs on thevm_diskbranch (if called_from_vm_disk:).So
vm_before.nicsis fresh and the new NIC resolves to a UUID, whilevm_before.disksis stale
and the new disk resolves to nothing and is dropped. Pass 2 succeeds because by then the disk
already exists on the VM.Worth flagging that in-code
# TODO are nics from vm_before used anywhere?— a maintainer was
unsure whether that refresh mattered. It does:set_boot_devices_orderconsumes it, and the disks
half is missing the equivalent.Two candidate fixes, both small: refresh
vm_before(or at least.disks) after_set_diskswhen
called from thevmmodule, so it mirrors the NIC path; or re-read the VM inside_set_boot_order
before resolving. Separately,set_boot_devices_ordersilently dropping a device the user
explicitly asked to boot from seems worth at least amodule.warn()regardless of how this is
fixed — that silence is what turns a resolution failure into a delayed boot failure.Relationship to #292 and PR #371
These are one boot-order cluster rather than three unrelated tickets:
-
🐞 Bug: vm_boot_devices likely does not correctly reboot VM #292 is the same stale-object family. Its TODO is still in the tree verbatim at
plugins/modules/vm_boot_devices.py:288:
# TODO BUG, this does not work - copy of 'vm' is used in say ensure_present.
Same root shape as this issue: avmobject that no longer reflects reality being used to act on
the VM. Fixing 🐞 Bug: Newly added disk/net device not added to boot_devices #289 and 🐞 Bug: vm_boot_devices likely does not correctly reboot VM #292 together is likely cheaper than separately, since both need the
refresh/ownership question answered once. -
vm_clone produces VMs that are silently fragile to multi-disk attach + reboot (no bootDevices, vda-anchored grub) #370 / PR 370 - vm_clone produces VMs that are silently fragile to multi-disk attach #371 touches the same shared code — it adds a caller of
set_boot_devicesin
plugins/module_utils/vm.py(+30/-1), with the fix being to set boot order after cloning and
report inmsgwhether it succeeded ("and boot order was set - you can change it with
vm_boot_devices module"). That explicitly-set-and-then-report pattern is close to what this issue
needs, so 370 - vm_clone produces VMs that are silently fragile to multi-disk attach #371 is worth reviewing with 🐞 Bug: Newly added disk/net device not added to boot_devices #289 in mind.To be clear about what I have not checked: I have not examined how PR 370 - vm_clone produces VMs that are silently fragile to multi-disk attach #371 constructs its cloned
VM object, so I make no claim about whether it is itself exposed to this staleness. That is a
question for whoever reviews it — and given 370 - vm_clone produces VMs that are silently fragile to multi-disk attach #371 has been open since June, a review pass looks
like the higher-value move either way.
Reproduction playbooks used for the above are available if useful.
-
Describe the bug
https://github.com/ScaleComputing/HyperCoreAnsibleCollection/tree/boot-order commit 0a080c7 exposed bug. Short description:
Check if this is still problem - it might be already fixed.