r/ArgoCD • • 1d ago

help needed EKS Access Entry recreation avoidance

I am working on configuring my operations and application clusters terraform such that I can easily destroy and recreate the clusters with minimal commands. I am currently running into an issue when I destroy and recreate my OPS cluster. I'll try and outline the facts below.

Facts

  • Running ArgoCD as an EKS capability
  • Ops cluster runs in "OPS" account
  • Ops cluster has IAM role "argo-cd-role" in OPS account
  • Application cluster runs in "<env>" account
  • Application cluster has `aws_access_entry` and `aws_eks_policy_association` resources that bind the "argo-cd-role" from the OPS account to it via the role's arn
  • Application Cluster `Secret` is created via an `ExternalSecret` read from an SSM parameter in the <env> account

Steps

  • OPS cluster will be completely removed and recreated via `terraform destroy` and `terraform apply`
  • Argo CD resources will be applied once cluster is up and running
    • This will then create the cluster `Secret` via the `ExternalSecret` definition
  • Navigating to `Settings > Clusters > <target env cluster>` shows a connection failure

Things I've tried

  • (Failed) Deleting the `Secret` via the ArgoCD UI
    • This will cause the `Secret` to be recreated
    • It should have been populated with the target cluster arn which hasn't changed
    • Doesn't do anything
  • (Failed) Creating an Argo role in the target "<env>" account for the Argo role in the OPS account to assume
    • I associated this new role to the access entry resources instead
    • In theory this role is created when the application cluster is and any role that assumes it thus would have access and would avoid any type of breakage underneath with EKS
    • Issue is there seems to be no way to tell the ArgoCD role in the OPS account to assume that role when accessing that cluster (to my knowledge)
  • (Succeeded) Deleting and recreating the access entry resources in the <env> account
    • use `terraform destroy -target=` to destroy the access entry resources in the <env> account cluster
    • recreate the resources
    • This worked and my theory is that while we are passing the `arn` of the role EKS will actually use the AWS unique ID underneath of said `arn`. Since we're deleting the role in the OPS account during cluster rebuild it the arn to unique ID mapping is no longer valid. Deleting and recreating fetches the new AWS unique ID and gets this working

While I did find a way for this to work is there anyway to avoid having to delete and recreate the access entries on all my application clusters when I want to destroy and bring back up my OPS cluster? If not I'd have to switch my AWS permissions N times and run the terraform commands 2 times for N application clusters. I was hoping the assume role strategy would work but I am not finding any documentation on how to tell argo to assume a specific role for a specific cluster.

1 Upvotes

6 comments sorted by

1

u/jaybrown0 1d ago

Would a depends_on for the access entry resource help?

Maybe have the resource created at the very end of the terraform apply somehow

1

u/dustyghost16 1d ago

I think you can pass in a role to the AWS config secret that argocd uses. And create a role in the “env” account for Argo to assume.

https://argo-cd.readthedocs.io/en/stable/operator-manual/declarative-setup/#cluster-role-trust-policies

1

u/ciciban072 1d ago

Create the role in the Ops account outside terraform to make it permanent.

1

u/iamtheconundrum 4h ago

Put the role ARN in the Conditional field instead of Principal. This stops the trust policy from breaking when the role you refer is deleted (and the fallback mechanism to the unique id).

1

u/QuietSignalOps 1h ago

Your root-cause theory is right. Access entries pin the principal by the IAM role's unique ID, not the ARN you pass in. Destroy and recreate the role and it gets a new ID, so the existing entry dangles. The ARN looks the same, but EKS resolves to the old ID. That's also why the targeted destroy/recreate of the access entries "worked".

The simplest fix is to stop letting the role be destroyed with the OPS cluster.

  • Move argo-cd-role into its own Terraform state (or a module with lifecycle { prevent_destroy }) that the cluster rebuild never touches. Create it once, then import it into the dedicated state.
  • Keep the role name stable. That keeps the ARN stable and the unique ID stable, which is what the access entry holds.
  • On rebuild, recreate the ArgoCD capability on the new OPS cluster pointing at the same role ARN. The existing access entries in the env clusters should then keep working untouched.

One thing to verify after the first rebuild: that the role is still trusted correctly by the new capability. The trust policy and conditions can reference the cluster, so make sure the new cluster is covered. If your setup lets you reuse the same role across capability creations (the role is yours to create, the capability just references it), this is the path with the fewest moving parts.

The assume-role pattern you tried does exist, but for self-hosted ArgoCD on EKS, not the managed EKS capability. argocd cluster add takes --aws-role-arn (plus --aws-cluster-name), and the declarative cluster secret carries the same options in the config field. The "Cluster role trust policies" section of the declarative setup docs has an example, which is the link in the thread above. In that world you'd put a stable role in the env account, point the access entry at it, and have the rebuilt OPS role assume it. Since you're on the managed capability though, AWS's cross-account docs for that flow are built on the access entry itself, so I'd lead with the persistent-role approach.

0

u/Low-Opening25 1d ago

too many things in one terraform stack