Situation
About 500 Windows edge kiosks were in the field on a major metro project. Contractors only saw a screen that would not work. Apps hung because the disk filled up: logs and data had nowhere to go, and nothing was cleaning up after itself. The customer was angry. The easy story was still “users” or “the supplier says the hardware is fine.”
Task
Stop treating five hundred identical failures like five hundred bad operators, and find the shared root cause before another site day burned on the same frozen screen.
Action
I ignored the blame loop and looked at the disk. Usable space was about 20 GB on devices that should have been 512 GB. Hardware leadership confirmed the estate’s real capacity; healthy units of the same model had it. On a sick box I dug in and found the truth: almost none of the drive had been partitioned for use. I fixed the layout, then wrote a SaltStack formula and pushed it across the fleet so we were not healing machines one remote session at a time.
Result
The fleet got the storage it already physically had. Full-disk hangs stopped being the default failure mode, the supplier argument lost its oxygen, and a metro project’s contractors got working screens again because someone questioned “20 GB on a 512 GB device” instead of closing another ticket.