Skip to content

🐞 Bug: Newly added disk/net device not added to boot_devices #289

Description

@justinc1

Describe the bug

https://github.com/ScaleComputing/HyperCoreAnsibleCollection/tree/boot-order commit 0a080c7 exposed bug. Short description:

# Bug:
# Task 1 - create VM with 0 disks, 0 NICs.
# Task 2 - add 1 disk, 1 NIC, make both bootable. Bug - only NIC was added to boot devices.
# Task 3 - same as task 2, nothing should change.

Check if this is still problem - it might be already fixed.

Activity

  1. self-assigned this
    on Dec 22, 2023
  2. ddemlow commented on Aug 10, 2026

    @ddemlow
    Member

    Re-validated against HyperCore 9.7.7.226383 (4-node cluster), collection main @ 34308d2,
    ansible-core 2.16.19. Answering the issue's own "Check if this is still problem - it might be
    already fixed"
    : still reproduces. Keeping open, with a sharper description and the mechanism
    below.

    Reproduction

    Ran the exact three-task shape from the description — create a VM with 0 disks and 0 NICs, then add
    1 disk + 1 NIC with both listed in boot_devices:

    TASK 1  create VM, 0 disks / 0 NICs   -> disks=0  nics=0  boot_devices=0
    TASK 2  add 1 disk + 1 NIC, both bootable
            -> boot_devices count = 1  (expected 2)
            -> the single entry is the NIC:
               {type: virtio, mac: 7C:4C:58:ED:D5:6F, vlan: 0, connected: true, ...}
            -> disks reports ['virtio_disk'] — the disk IS attached, just absent from boot order
    

    So the original report is exactly right: only the NIC lands in boot order.

    It self-corrects on a second pass — which makes it easy to miss

    Running the identical task again converges:

    PASS 1  changed=True   boot_devices=['virtio']
    PASS 2  changed=True   boot_devices=['virtio_disk', 'virtio']
    PASS 3  changed=False  boot_devices=['virtio_disk', 'virtio']
    

    This is the same two-pass shape as #30's original report. It still matters, because anyone who
    runs their playbook once gets a VM whose disk is silently not bootable
    — no error, no warning,
    and changed=True reports success. For a VM that is about to be booted from that disk, the failure
    surfaces later and far from its cause.

    Workaround: vm_boot_devices with state: set fixes it immediately afterwards —
    boot_devices becomes ['virtio_disk', 'virtio']. So the boot-order write path itself is fine;
    the problem is what the vm module feeds into it.

    Mechanism (code-verified on main@34308d2)

    Three things compose:

    1. The VM object used for boot order is captured too early. ensure_present
    (plugins/modules/vm.py:448-460) reads vm_before, then runs _set_disks and _set_nics, and
    only then calls _set_boot_order(module, rest_client, vm_before, existing_boot_order) — passing the
    object it captured before the devices were created.

    2. Unresolvable boot devices are dropped silently. set_boot_devices_order
    (plugins/module_utils/vm.py:769-776):

    def set_boot_devices_order(self, boot_items):
        boot_order = []
        for desired_boot_device in boot_items:
            vm_device = self.get_vm_device(desired_boot_device)
            if not vm_device:
                continue                      # <-- silent drop, no error or warning
            boot_order.append(vm_device["uuid"])
        return boot_order

    3. This explains why the NIC survives and the disk does not. The NIC path does refresh the
    object it is about to hand on (plugins/module_utils/vm.py:1606-1610):

    updated_virtual_machine_TEMP = VM.get_by_old_or_new_name(module.params, rest_client=rest_client)
    updated_virtual_machine = vm_before
    updated_virtual_machine.nics = updated_virtual_machine_TEMP.nics
    del updated_virtual_machine_TEMP
    # TODO are nics from vm_before used anywhere?

    The disk path does not. When ManageVMDisks.ensure_present_or_set is called from the vm module it
    takes the return changed path at vm.py:1450; the get_vm_by_name refresh at vm.py:1443 only
    runs on the vm_disk branch (if called_from_vm_disk:).

    So vm_before.nics is fresh and the new NIC resolves to a UUID, while vm_before.disks is stale
    and the new disk resolves to nothing and is dropped. Pass 2 succeeds because by then the disk
    already exists on the VM.

    Worth flagging that in-code # TODO are nics from vm_before used anywhere? — a maintainer was
    unsure whether that refresh mattered. It does: set_boot_devices_order consumes it, and the disks
    half is missing the equivalent.

    Two candidate fixes, both small: refresh vm_before (or at least .disks) after _set_disks when
    called from the vm module, so it mirrors the NIC path; or re-read the VM inside _set_boot_order
    before resolving. Separately, set_boot_devices_order silently dropping a device the user
    explicitly asked to boot from seems worth at least a module.warn() regardless of how this is
    fixed — that silence is what turns a resolution failure into a delayed boot failure.

    Relationship to #292 and PR #371

    These are one boot-order cluster rather than three unrelated tickets:

    Reproduction playbooks used for the above are available if useful.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions