Situation
Saturday, 7am. Guests with valid passes were stuck outside the gates while the campus was filling up for a busy weekend. The university’s housing and guest systems, software and hardware we support, had stopped letting people through. They’d reached out even though I wasn’t working and wasn’t on call.
Task
Get guests moving again fast, keep other enterprise edge sites from walking into the same failure, and stay the single voice for a customer already under pressure. I’d been on site during quieter setup days, so I could picture the queues, the staff stress, and visitors who’d done everything right still locked out. That landed. I messaged my manager, flagged it as a direct P1, and they trusted me to lead it.
Action
I owned the customer thread so they weren’t bouncing between teams, and ran two internal tracks in parallel. Aging hardware got a temporary fix to restore access, then account management was engaged for replacements. Deeper down, a missing server wasn’t just this campus’s problem: other enterprise customers with denser edge devices could have hit the same wall. I brought in major incident management, walked them through a product they didn’t live in day to day, and we chased that root cause until it closed.
Result
Guests were getting in again, hardware had a real replacement path, and the infra gap wasn’t left waiting for the next busy weekend somewhere else. The part that stuck wasn’t the severity tag. It was choosing to own someone else’s chaotic morning when I could have said I wasn’t rostered.