| title | OpenStack Ironic Node Stuck in Deploying | ||||||
|---|---|---|---|---|---|---|---|
| slug | openstack-ironic-node-stuck-in-deploying | ||||||
| technologies |
|
||||||
| severity | high | ||||||
| tags |
|
||||||
| related |
|
||||||
| last_reviewed | 2026-06-27 |
$ openstack baremetal node list
+--------------------------------------+----------+---------------+-------------+
| UUID | Name | Provision | Power State |
| | | State | |
+--------------------------------------+----------+---------------+-------------+
| 9c1f... | node-07 | deploying | power on |
+--------------------------------------+----------+---------------+-------------+
ironic-conductor[3120]: ERROR ironic.drivers.modules.agent_base [req-...] \
Deploy failed for node 9c1f...: Timeout reached while waiting for the IPA ramdisk \
to start. last_error: "Timeout reached while waiting for callback from ramdisk"
A bare-metal node is wedged in the deploying (or wait call-back) provision
state and never reaches active. Ironic deploys by powering the node into a PXE/
iPXE-booted IPA (Ironic Python Agent) ramdisk, which then writes the image to disk
and calls back to the conductor. If any step in that chain stalls β PXE/DHCP,
ramdisk boot, network reachability back to the conductor, or the disk write β the
node stops progressing and eventually times out or hangs.
- openstack (ironic-conductor, ironic-python-agent ramdisk, neutron/PXE-DHCP, TFTP/HTTP)
high β the target node cannot be provisioned and is held out of the pool. If the cause is shared infrastructure (DHCP, TFTP, the provisioning network), many deploys fail at once.
- PXE/DHCP failure β the node never gets an address or boot file on the provisioning network (wrong VLAN, DHCP not serving, TFTP/HTTP unreachable).
- The IPA ramdisk cannot route back to the conductor's callback URL (firewall,
wrong
[deploy] http_url/api_url, MTU mismatch). - Wrong NIC/boot mode β node boots from disk instead of network, or BIOS vs UEFI
mismatch with the configured
boot_mode. - BMC/IPMI problems so the conductor cannot reliably control power.
- Disk/RAID issue: no valid
root_devicehint match, or the target disk fails. - A conductor restart/crash leaves the node's lock and TaskManager state stale.
Ironic's deploy is a multi-stage state machine: set boot device to PXE, power on,
serve DHCP + boot file, boot IPA, IPA pulls the image over HTTP and writes it,
then IPA POSTs a callback (heartbeat) to the conductor, which finalizes and
reboots into the new OS. The node sits in deploying/wait call-back until the
heartbeat arrives. A break anywhere β L2 reachability, boot config, BMC control,
or disk β stops the heartbeat, so the node never advances and times out. The
conductor log names the stage that stalled.
# Provision state and the last error
openstack baremetal node show node-07 -c provision_state -c last_error -c power_state -f value
# Conductor log for this node's deploy
journalctl -u ironic-conductor --since "30 min ago" | grep -i 9c1f
# Is the node's power controllable via the BMC?
openstack baremetal node power show node-07 # or: ipmitool -I lanplus -H <bmc> ... power status
# Validate the node's driver/interfaces are all OK
openstack baremetal node validate node-07
# Is DHCP/PXE serving on the provisioning network?
journalctl -u neutron-dhcp-agent --since "30 min ago" | grep -iE "DHCPOFFER|DHCPACK"
ss -lunp | grep -E ':67|:69' # dnsmasq DHCP/TFTP listening
# Confirm the deploy/agent images and ramdisk URLs resolve
curl -sI http://<conductor>:8088/agent.kernel | head -n1$ openstack baremetal node show node-07 -c provision_state -c last_error -f value
deploying
Timeout reached while waiting for callback from ramdisk
$ openstack baremetal node validate node-07
+------------+--------+----------------------------------+
| Interface | Result | Reason |
+------------+--------+----------------------------------+
| power | False | IPMI call failed: power status. | <-- BMC issue
| deploy | True | |
+------------+--------+----------------------------------+
# Healthy: validate is all True, DHCPACK appears for the node's MAC, and the
# provision state advances deploying -> wait call-back -> active.
- Read
last_errorandnode validateto pin the failing stage, then fix that stage specifically:- PXE/DHCP: correct the provisioning VLAN/network, ensure dnsmasq serves DHCP and TFTP/HTTP, verify the node's port MAC is registered in Ironic.
- Callback/network: open conductor
api_url/http_urlports to the provisioning subnet; fix MTU. - Boot mode: align
boot_mode(uefi/bios) and set PXE first in BMC. - BMC: fix IPMI/Redfish credentials/connectivity so power control works.
- Disk: set a correct
root_devicehint or replace the failed disk.
- Once infrastructure is fixed, unstick the node β abort and clean back to
available, then redeploy:openstack baremetal node abort node-07 # if abortable openstack baremetal node maintenance set node-07 --reason "stuck deploy" openstack baremetal node maintenance unset node-07 openstack baremetal node deploy node-07
- If a conductor crash left a stale lock, restart
ironic-conductorso the node is re-managed.
openstack baremetal node show node-07 -c provision_state -f value # Expect: active
openstack baremetal node validate node-07 # all True- Monitor the provisioning network (DHCP/TFTP/HTTP) and conductor reachability.
- Standardize BMC firmware, boot mode, and NIC PXE order via inspection/templates.
- Set
[conductor] deploy_callback_timeoutsensibly and alert on nodes indeployingpast it. - Run periodic
node validateacross the fleet to catch BMC drift early.
openstack Β· ironic Β· baremetal Β· deploying Β· pxe Β· production