This was found on a racklette after sending heavy IO to ~190 instances then rebooting a seld.
Things eventually finished except for one downstairs that appears to have gotten stuck trying to re-open an extent after a repair.
We see that a propolis server has a downstairs that is in LR:
oxz_propolis-server_a12fa990 11608 440985a8 6ccc7d10 ACT ACT LR
The propolis server just reports that it has work to do and is not making progress:
BRM42220081 # tail -f $(oxlog logs --current oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8) | looker
17:26:41.607Z WARN propolis-server (vm_state_driver): flush check fired despite having jobs; resetting it
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
17:26:42.608Z WARN propolis-server (vm_state_driver): flush check fired despite having jobs; resetting it
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
17:26:43.258Z INFO propolis-server (vm_state_driver): request completed
file = /home/build/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/dropshot-0.16.7/src/server.rs:867
latency_us = 2166
local_addr = [fdd5:cfdb:cca5:104::1:356]:55639
method = GET
remote_addr = [fdd5:cfdb:cca5:101::5]:47691
req_id = 710c4be1-679d-4092-beb0-a46bec2d7e36
response_code = 200
uri = /08589f2b-eb0f-4110-a3e8-f1a54e861f54
17:26:43.609Z WARN propolis-server (vm_state_driver): flush check fired despite having jobs; resetting it
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
17:26:44.610Z WARN propolis-server (vm_state_driver): flush check fired despite having jobs; resetting it
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
17:26:45.611Z WARN propolis-server (vm_state_driver): flush check fired despite having jobs; resetting it
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
17:26:46.612Z WARN propolis-server (vm_state_driver): flush check fired despite having jobs; resetting it
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
Going back in time a bit, I can find the propolis logs that started the repair:
2026-04-14 05:11:54.664Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: Returning UUID:eb39f795-fee1-45fb-a509-96eb25b98aba matches
= downstairs
client = 2
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
2026-04-14 05:11:54.771Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: client 2 is ready for live-repair
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
2026-04-14 05:11:54.771Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: Create new job ids for 0: ExtentRepairIDs { close_id: JobId(15483549), repair_id:
JobId(15483550), noop_id: JobId(15483551), reopen_id: JobId(15483552) }
= downstairs
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
2026-04-14 05:11:54.771Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: RE:0 repair extent with ids 15483549,15483550,15483551,15483552 deps:[[JobId(15483
547)]]
= downstairs
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
2026-04-14 05:11:54.771Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: 15483552 final dependency list [[JobId(15483551)]]
= downstairs
client = 2
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
2026-04-14 05:11:54.771Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: 15483549 final dependency list [[]]
= downstairs
client = 2
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
2026-04-14 05:11:54.771Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: started repair ffa54de6-cd50-4a3d-bc56-01f10046af08
= downstairs
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
2026-04-14 05:11:54.771Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: new DNS resolver
addresses = [[fdd5:cfdb:cca5:1::1]:53, [fdd5:cfdb:cca5:2::1]:53, [fdd5:cfdb:cca5:3::1]:53]
job = notify
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
2026-04-14 05:11:54.891Z INFO propolis-server/11608 (vm_state_driver) on oxz_propolis-server_a12fa990-1cc5-4593-898f-018bf7178bf8: notified Nexus of live repair start
job = notify
session_id = 6ccc7d10-3083-444a-a8a9-7b0f9a0617db
The last logs I have from the downstairs show it was doing a LiveRepair:
2026-04-13 17:44:55.753Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: eid:45 Found repair files: ["02D"]
2026-04-13 17:44:57.554Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: Verify extent 45 still ready for copy
2026-04-13 17:44:57.555Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: 1 repair files downloaded, move directory "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/02D.copy" to "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98ab
a/00/000/02D.replace"
2026-04-13 17:44:57.555Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: Copy files from "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/02D.replace" in "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000"
2026-04-13 17:44:57.685Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: Move directory "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/02D.replace" to "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/02D.completed"
2026-04-13 17:44:59.425Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: Created copy dir "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/005.copy"
2026-04-13 17:44:59.425Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: eid:5 Found repair files: ["005"]
2026-04-13 17:45:02.162Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: Verify extent 5 still ready for copy
2026-04-13 17:45:02.164Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: 1 repair files downloaded, move directory "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/005.copy" to "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98ab
a/00/000/005.replace"
2026-04-13 17:45:02.165Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: Copy files from "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/005.replace" in "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000"
2026-04-13 17:45:02.442Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: Move directory "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/005.replace" to "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/005.completed"
2026-04-13 17:45:02.451Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: Created copy dir "/data/regions/eb39f795-fee1-45fb-a509-96eb25b98aba/00/000/03C.copy"
2026-04-13 17:45:02.452Z INFO crucible/15946 on oxz_crucible_47c99cd0-13ae-4f31-a08a-2ca1a1179a78: eid:60 Found repair files: ["03C"]
But, no more output has come from the downstairs since then.
If I hit the repair endpoint on the downstairs with /work, we get a little bit of info:
17:33:02.729Z INFO crucible: accepted connection
local_addr = [fdd5:cfdb:cca5:102::8]:23002
remote_addr = [fdd5:cfdb:cca5:104::1]:63441
task = repair
17:33:02.730Z INFO crucible: request completed
latency_us = 70
local_addr = [fdd5:cfdb:cca5:102::8]:23002
method = GET
remote_addr = [fdd5:cfdb:cca5:104::1]:63441
req_id = 47013534-9246-45d0-818e-b975a7e7c92f
response_code = 200
task = repair
uri = /work
17:33:02.732Z INFO crucible: Active Upstairs connections: [ConnectionId(86)]
17:33:02.732Z INFO crucible: Crucible Downstairs work queue:
17:33:02.732Z INFO crucible: JOB_ID IO_TYPE STATE DEPS
17:33:02.732Z INFO crucible: 15483552 ReOpen [JobId(15483551)]
17:33:02.732Z INFO crucible: Completed work [JobId(15483549)]
So the downstairs is still "alive" in some sense. And the upstairs has not kicked it out due to timeouts, so it's still responding to pings from the upstairs.
This was found on a racklette after sending heavy IO to ~190 instances then rebooting a seld.
Things eventually finished except for one downstairs that appears to have gotten stuck trying to re-open an extent after a repair.
We see that a propolis server has a downstairs that is in LR:
The propolis server just reports that it has work to do and is not making progress:
Going back in time a bit, I can find the propolis logs that started the repair:
The last logs I have from the downstairs show it was doing a LiveRepair:
But, no more output has come from the downstairs since then.
If I hit the repair endpoint on the downstairs with /work, we get a little bit of info:
So the downstairs is still "alive" in some sense. And the upstairs has not kicked it out due to timeouts, so it's still responding to pings from the upstairs.