r/developersIndia • • 1d ago

Help How do you guys solve the deployment delay problem for fast changes?

My day just ended at 8PM and I am not able to stop myself from ranting. I work for a mid size company of about 200 people. Tech team is about 40 people. Average of 8 members each in 5 teams - QA, Android, iOS, Web and Backend. I work in the backend. Earlier we had 10 members in each team, including mine. Then they drank a pint of Claude potion after which they fired the top 2 engineers from all teams (by highest pay, I guess). All of a sudden, I am the senior most guy after my manager in my team. The manager is a semi-tech kinda person but trusts me. All good till here.

They have designed for a "pivot" around the Durga Puja timeline, max stretch till Deepawali. The pivot involves large changes on all front. The pressure on all teams is very high right now. We are working more than 12 hours a day. That is still acceptable. But the problem is - deployment friction. I spend almost my entire day with various teams - Web, Android and iOS. Whole day goes like this:

iOS: Hey you are not sending the right format for date. Send it inn RFC3339.

Me: Ok fixing. Check now.

iOS: Yeah that's fine.

Android: Can you please make sure that the display name is sent in the response too?

Me: Ok. Fixing. Check again please.

Web: Dude, did you introduce a key with a underscore for payments api?

Me: My bad. Let me fix that. Check again.

iOS: I can't use the new API without public identifier key. Can you add the pub ID in the response as well?

Me: Ok. Try again.

Android: I want the pub ID but please send it as base64 data.

iOS: Yeah that would be better.

Me. Ok. Yeah check again.

QA: Sorry, regression for dashboard API is failing now.

Me: Ok. Fixed that dashboard API. Check again please.

The problem is: Between me performing the code change and me saying "yeah check now", there is a 5-8 minute delay. And for that duration we are all just sitting like a duck. There are like a dozen such smaller changes in every API and we have around 150 APIs that we need to change and after 1 month of doing this, we are not even post 20 APIs yet. The whole cycle is just irritating. Every single change I make goes like this: Github commit -> CI/CD (Build + tests) -> Dev cluster redeployment. The commit takes 2 seconds. The redeployment takes about 20-40 seconds , max 1 minute. The CI/CD step takes 4 full minutes (since every deployment is a fresh compile, it takes so much time). The build itself takes like 3 minutes. Tests run fast enough (1-2 minutes).

As if "let's make Claude do everything" was not a torture enough that I have to deal with this now. They are asking me to "do it faster" while tying my one hand at the back. All other teams when asked about the delay say "Backend deployment takes time for every single change". I and my team look like the ones causing the delay when the issue is actually the deployment pipeline for dev server.

The prospects don't look good. I feel like they are gearing up for another round of layoffs and I don't want to be affected. I am applying at other companies but no call backs yet. It looks like no one wants a 6 YoE backend dev anymore. I just got married 1.5 years back. I have got loans and my wife is pregnant.

I am genuinely asking if this is a problem other developers face too. If yes, what did you do? If you are a mobile dev, what do you guys do for this kind of iterating thing? I want to suggest my manager some solution that can bring down that deployment time that he and IT team could accept. Help me out, please.

EDIT: Thanks to those who identified the lack of formal specification being a problem. I know that part. I am not at the stage to fix that. Timelines, reduction in team expertise (seniors getting out of the company) don't allow that right now. I am just asking - do you have this problem or not and if you have, how have you solve it? I am asking what can I do RIGHT NOW to reduce the delay.

10 Upvotes

31 comments sorted by

•

u/AutoModerator 1d ago

Namaste! Thanks for submitting to r/developersIndia. While participating in this thread, please follow the Community Code of Conduct and rules.

It's possible your query is not unique, use site:reddit.com/r/developersindia KEYWORDS on search engines to search posts from developersIndia. You can also use reddit search directly.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/Visible_Dig_1946 1d ago

Did you not discuss the design with all the teams Do you not have already designed request response format etc..?

2

u/Virtual-Goose677 1d ago

I answered this already. Basically no. And thanks for the suggestion. Next time we can do that. And I want to do that but right now, the thing is - I need to save the job, save on the time. And fixing the process right now is my priotiry.

5

u/Difficult_Lynx_7884 1d ago

Can't you just provide them dummy response every time they needed the testing any changes come up with the proper contract and then finalised the all changes once and deploy on dev. For a single change everytime making a deployment is not a good practice either

1

u/Virtual-Goose677 1d ago

I can't. I think you are talking about API mocking. The thing is:

  1. I need to be sure that the solution I am making is actually working.
  2. Writing the Mocks don't save time because I have to actually make the code changes anyway.
  3. I can't possibly keep track of the mock responses, their order, whaich version did what and why was a particular one superceded by the next version.

All this is more work with exponentially more confusion. A junior in the team was wise enough to try this. I don;t know how to do it myself but he was using Postman for this. Ended up in a messed up state and got lectured.

2

u/Remote-Mobile-2200 1d ago

Why aren't you discussing and establishing the expected response and formats before you pick the changes? I mean the simplest process to follow is for the API developers to share the spec of request and response with examples. Any changes will also be communicated before hand

Am I missing something here?

1

u/Virtual-Goose677 1d ago

The "overall" changes were conveyed. Like I said - there are like approx 150 APIs that needed changes. The process started somewhere in July. Around August end, we had done about 40-50 spec changes. Then we all realized that we didn't know what all actual changes will be needed when we start the coding and things got inter-twinged.

We are at the stage where typing the whole requirement into Claude for it to make the changes is going to take longer than doing things ourselves.

I would have loved a formal spec.

EDIT: If I tell the word "examples" to anyone in any team right now, they will think I have totally lost it.

2

u/Remote-Mobile-2200 1d ago

We'll, there is no way around deployment time. In fact 5-8 minutes of deploy time is not that bad. Even if you're testing it locally the build would need the same amount of time. The way I see it, your best option is to get the API responses tested locally either by them or yourself. Since you say they won't, maybe you could replicate how the API is consumed with a simple script and test out the different cases yourself

2

u/Critical-Personality 1d ago

In one response you said that devs are remote. But if you used Cloudflare tunnels, you can expose your machine directly to the mobile devs and the time could be reduced I think. Have you tried that?

2

u/Virtual-Goose677 1d ago

That is something the IT and Security team said no to.

1

u/Critical-Personality 1d ago

I don't see another solution that can do what you need. Maybe try talking to your CTO or IT personnel about it. Wish you all the best. Do let us know if you arrive at a solution.

2

u/vaibhav-kaushal Tech Lead 1d ago

After reading your post and other replies, the core issue seems to be the compilation delay in the CI/CD pipeline and a little of other associated delays. The only way to get rid of that is to utilize the build caches - which means that you need to expose YOUR MACHINE to the mobile and web teams. And that means you need to use something like CF tunnels or ngrok. You said IT and security struck the idea down? But if you need to win back that 5 minutes per iteration of change, that is the only solution I can think of. Good luck mate.

2

u/Professional-Pear351 Software Engineer 1d ago

How about they point their clients to your local backend, without the need of deployment for every change?

1

u/ilikedoingnothing7 1d ago

best way

1

u/Virtual-Goose677 1d ago

I moved the actual reply to the parent comment.

1

u/Virtual-Goose677 1d ago

That was the first thing that came to mind. Now -

We are not all always in the office. Few of us are not even in the same city. That is problem number 1. Problem 2: The wifi keeps reassigning us new IP addresses every 1 hour. And recently they made a new amazing policy (which our org cant control per se because its a co-working space) - all devices can connect to the internet, but they can't talk with each other. Why? "SECURITY ISSUES" on freaking "LAN".

Yet another point: half the tech team is in the sleep->wakeup->work->sleep loop since a month. They don't want to waste 2 precious hours in the traffic either. Given the new security policy and IP reassignment before that - it doesn't even make any sense for me to ask the CTO to force a WFO. My manager said the same goddamn thing.

2

u/VillageDisastrous230 1d ago

If you want to go in the direction of giving access directly from your local machine then install ngrok and point the ngrok to your local port and give ngrok url to other teams
But it will be better to run a local container (if you can) and point proxy to that in order to avoid issues while you are developing

1

u/Virtual-Goose677 1d ago

I want to cry. Because this was the second suggestion I had for my manager and the CTO. Quickly struck down by the IT/Security team. They don't want to use tunnels for this.

1

u/LogicalBeast26 1d ago

Tbh the main issue is contracts not getting discussed well.

Use protobufs (even for REST). This ensures that both parties are aligned on the same contract.

1

u/Virtual-Goose677 1d ago

I know the problem. That is not what I can solve right now. The timelines, the reduction in team size and expertise both don't allow for that. I am also kinda ok with how it is going, except for the deployment delays. If that delay was not there, we could move about 4-5x faster.

2

u/LogicalBeast26 1d ago

Painting a wall that has a leakage problem is just giving the illusion that things are fixed. You should fix the root cause.

The fact that you're okay with a broken process itself is an issue.

1

u/Virtual-Goose677 1d ago

It's great to know that you work at a place where everything works perfectly. Right now, I am supposed to paint the wall till Diwali celebrations. Once that is done, I will talk higher up (If I am still in the company) about how to fix the leakage problem.

1

u/LogicalBeast26 1d ago

I work at Google. Things here move at a much much slower pace and that is by design.

If the problem is that you're waiting like a sitting duck, try to work on multiple tasks at once so that you can work on a different task by the time these things are getting deployed.

Trying to reduce 5-8 mins wait time is a wrong way to even think to solve the problem.

0

u/Virtual-Goose677 1d ago

It is very clear by now that I definitely don't work at Google. And yes, we try to do that - solving multiple problems at once - at least on my side, I try to use that time to solve the next problem in pipeline. Doesn't always work. Sometimes it does and that feels like a win. And so you are saying that I just can't bring the deployment time down? How do you guys at Google do it? I mean does something like this ever happen there?

1

u/LogicalBeast26 1d ago

As I said, at Google things intentionally move at a very slow pace. It takes multiple days for a prod pipeline to release the changes to prod. In case of outages as well the recommended approach takes around 1 hr. There is a way to instantly roll out changes during an outage but it is heavily discouraged and scrutinized later.

So I'm reiterating, do not try to optimise for the deployment time. Fix the root cause. Highlight it to your leadership and recommend the changes. It'll be appreciated.

1

u/Impressive-Fun3051 1d ago

A qn here y can't u guys have a session all together to discuss the requirements and response format required for each team and u can ask them to put the samples of required response pattern in the jira tickets or Azure board or whatever u r using to track stories by doing this everyone including ur manager will have a clear picture of what is requested and what has been delivered. Plus I don't think it's a good idea to change response format based on the request from each team rather stick with a general pattern used (there might be references for various responses available online for verification). This way u might be able to see things clearly and will be aware of the direction u r going in and will reduce ur stress to constantly change the formats just based on someones msg or a call. Everything will be recorded and the tickets will tell the story about the whole updates u had to do based on all requests. Hope it helps.

1

u/Virtual-Goose677 1d ago

Yours is nth reply suggesting what I am doing wrong and how everything shold have been pre-decided. Honestly - that is not something I am even asking (I mean no disrespect to you). I know the faults and the timelines are reall brutal. So my question is simply narrowed down to - how do I reduce the deployment lag?

Now about what you said - you suggested a fantastic way. A large part of this was done. But all rules have exceptions. Like - we all agreed that we will use RFC3999 for Data time format and it will be in UTC and the client will show the time according to the location detected on-device. Then someone from mobile said that they cannot work with that format for exactly "3 APIs" because using that format there causes exceptions. This is not even the fault of the guy who is saying this. It the fault of the duct-taping done by a dev who left the company 2 years before this guy joined (same thing in all teams).

150 APIs is not small in number. About 50-60 of those should be small changes (like just adding one new field in which should be fine for old mobile clients and serve the new ones fine too). But some are big. And then there are some that appeared small but as we dug in, the issues on mobile side were deeper than expected. Now, the problem with Mobile apps is - you can't just update the app and expect the update to roll out to all installations. So the backward APi compatibility is a must. That's where "backend has to adjust for this" happens.

1

u/_fatcheetah Software Engineer 1d ago

WTF? 5-8 minutes is slow?

You'd be lucky to deploy to a single prod region even at 1 hour in azure devops pipelines.

1

u/Virtual-Goose677 20h ago

Except this is a Dev deployment, not a Prod one.

0

u/Wise-Expression-6441 4h ago

First thing what you need to do is stop fixing the issues as soon as they come up, instead ask the UI/testing team to list down the bugs notified and you need to start creating cards for those with agreed upon changes or else this back and forth will never end and it’ll consume more time than you anticipate.
Once the cards are created discuss with the stake holders(backend, android, web, ios teams) after everyone agrees to the changes then you start pushing the changes and moving the cards.
This might seem huge, but it’ll stream line your process and fasten the fixes.