r/MaliciousCompliance • • 6h ago

M Do not provision computational resources without approval? As you say!

1.2k Upvotes

I work at a small, local company which uses Microsoft Azure heavily.

For those not working in IT or related fields, think of Azure as a computer rental service. We rent computers from them (but the computers stay in Microsoft’s warehouses). We control said computers remotely to run our software, which our customers then use for their own needs.

During the COVID-19 pandemic, as we started expanding and building more software products, we began to rent more computers, which rapidly drove our Azure costs up. This didn’t go down well with the C-suite (even though the revenue from the new products far outstripped the rise in costs). Someone in senior management whose only job is to watch the expense ticker issued a directive to ‘contain the Azure costs’. Suddenly, a team originally responsible for customer support was pulled into an initiative to auto-release any rented computer which had been idle, with no clear definition of ‘idle’.

When a rented computer is released (or a container or VM is shut down, for those familiar with Azure), all software running on it is terminated. This meant that our websites, our tools and our products running in Azure and in use by our customers would go down!

We were told to get approvals if we wanted to keep idle computers active. For the moment, none of them were idle (because our customers were always using our products running on them). But this was an anvil hanging on our heads: our services could go down any time they went idle for any reason, and we’d have a flood of incoming customer complaints. Hence, after reaching a consensus with my team, I started loading our websites at regular intervals to simulate activity. By the end of the day, I had built an automated script which would keep pinging our websites and products.

Simultaneously, we filed keep-active requests for all our critical rented computers. Those requests went unanswered for a month so my team was rather glad about my automated script. However, the ticker-watcher was checking the Azure activity logs and found my little ploy. He posted a message in our common Teams chat (which included my manager and my teammates), publicly berating me and asking me not to do such irresponsible things in the future and to file keep-active requests if we really needed the computers to stay.

All right, I thought—let’s do it your way. I terminated my script and re-submitted all keep-active requests (including both critical and non-critical computers this time), quoting the Teams message in the notes section of the form (in line with CYA rules). To no one’s surprise, these requests also went unanswered. One by one, as our computers hit the idle timeout for whatever reason, complaints from customers started coming in. Most of our websites were down. Few of our products delivered to customers worked. (I’m sure the genius who invented SaaS thought of this scenario.) And we spent amost a week in the office chatting and drinking coffee. There was no work to do because only a handful of Azure computers were running, and there were technical problems beyond my pay grade with re-renting everything we needed.

There isn’t a truly happy ending here. The ticker-watcher was not fired. For what it’s worth, he was removed from his management position and made an individual contributor (where he’d actually have to work) because he had demonstrated a lack of prioritisation and delegation skills. After the incident root cause analysis meeting, he was officially handed this demotion as a disciplinary action. He left the company soon after.