GitOps Reconciliation Drift With Argo CD and Application Sets
ApplicationSet's separate reconciliation clock creates undetectable drift across entire fleets.

Drift in Argo CD is a structural property of any system where a desired state and a live state get maintained by two separate processes running on two separate clocks. The Application Controller runs a continuous comparison loop, checking live cluster state against the target state derived from Git, and marks any divergence as OutOfSync. Two distinct sources of drift surface identically in the UI: a new Git revision that changed the target state, or an out-of-band actor that changed the live resource, both produce the exact same symptom but call for entirely different remediation. One means the cluster needs to catch up to a legitimate change. The other might mean someone made an emergency fix that Git doesn't know about yet.
The default reconciliation interval is 120 seconds, and Git webhooks compress that window to near-immediate. Neither approach eliminates the interval during which drift can accumulate undetected. SyncPolicy options extend the machinery around this loop: selfHeal reverts live drift even when Git hasn't changed, automated sync applies new Git commits without manual intervention, and prune removes resources no longer declared in Git. None of these mechanisms can tell the difference between an accidental change and a sanctioned emergency one; auto-sync reconciles state, but it has no way of knowing why the state diverged.
Three patterns account for most single-application drift in practice. Manual kubectl edits are the obvious case. Less obvious are mutating admission webhooks that inject fields absent from the Git manifest, which can flip an application straight to OutOfSync moments after a clean sync completes. The canonical case is an HPA adjusting spec.replicas while Argo CD continuously resets the field back to whatever the manifest says, producing a revert loop that never settles. That loop reappears in a much more damaging form once selfHeal gets applied across a fleet rather than a single application.
The ApplicationSet controller's second, slower clock
Everything above describes one application, one controller, one reconciliation loop. ApplicationSets change that picture by adding a second controller that runs on its own, slower cadence, and that timing gap has no equivalent in single-application Argo CD. The Application Controller reconciles every 120 seconds. The ApplicationSet controller defaults to a three-minute interval, so generator re-evaluation always trails behind the loop that actually syncs resources to the cluster.
The division of labor matters here. The ApplicationSet controller reconciles ApplicationSet resources and produces Application objects from them; the Application Controller then picks up those generated Applications and reconciles them against the cluster. Two controllers, two clocks, one desired state that both are supposed to agree on. A cluster selector generator or a Git directory generator may not detect a new match, or a removed match, for up to three minutes, and during that window the fleet's actual membership and its declared membership diverge.
There's a sharper trap hiding in this timing gap. Editing a generated Application directly, as an emergency measure, feels like the fastest fix, but it isn't reliable: the owning ApplicationSet will silently overwrite those fields on its next reconcile. The correct intervention point is the ApplicationSet itself, or whatever temporary-toggle mechanism it documents. Anyone who has tried to patch a generated Application under pressure and watched the patch vanish three minutes later has learned this the hard way.
Visibility into this layer has historically been thin. Before Argo CD's v3.2 release, the ApplicationSet controller could update its status silently, so operators had no clear record of what changed or when, and that release improved visibility into how the controller tracks and reports its own status changes. That improvement matters precisely because the ApplicationSet layer is where drift is hardest to see coming.
Fleet-wide generator patterns create new drift surface beyond per-application
Once a generator defines fleet membership, a bug in that generator is not a contained, single-application problem. It propagates to every Application the generator produces, and if it produces consistent drift across the fleet, that drift can look enough like intended state to escape notice. Cluster generators, Git directory generators, and matrix generators each carry their own re-evaluation latency and their own blast radius when something in the generator logic is wrong.
Ordinary environment duplication produces the most common version of this. Teams frequently run near-identical ApplicationSets for staging and production that differ only in a path parameter. A path change made in one and not mirrored in the other causes the two environments to diverge without triggering any alert, since from each generator's perspective it is doing exactly what it was told to do. This has been noted as an accepted trade-off for environment isolation, but it is a trade-off, and it should be treated as one rather than assumed away.
ArgoCD's ApplicationSet custom resource can exceed the 256 KB metadata annotation limit on client-side apply once a fleet grows large enough. That failure appears only past a certain scale and then stops the apply outright.
Underneath both of these is a consolidation risk that rarely gets named until it causes an incident. When staging and production ApplicationSets are maintained separately and diverge in ways that are not immediately visible, the fleet's declared state and its live state can diverge across environments without any individual application appearing OutOfSync. The failure is not visible at the application level because it was never designed to be checked at that level. Fleet-wide drift also erodes the operational record of what's actually deployed and when it last changed. GitOps promises an audit trail, and that promise weakens every time generator logic introduces variation nobody documented.
The controller-conflict failure modes that selfHeal makes worse at scale
SelfHeal does not remove drift at scale. Applied without scoping, it turns ordinary controller activity into a continuous revert loop that eats reconciliation capacity and hides the problems operators actually need to see. The HPA case from earlier becomes the clearest illustration once selfHeal is switched on fleet-wide: the HPA adjusts spec.replicas, Argo CD resets it to the manifest value, and the pod count churns on every single reconciliation pass, with the application flipping endlessly between Synced and OutOfSync.
Secrets carry their own version of this fight. When an external secrets operator updates values, or a Vault sidecar rotates credentials, or some other controller mutates a secret after creation, Argo CD reports the change as drift. SelfHeal reverts the rotation, and the external system dutifully re-applies it a moment later.
The Argo Rollouts case is the sharpest example of two reconcilers fighting over the same resource. Argo Rollouts updates traffic-routing objects, Istio VirtualService weights, NGINX Ingress rules, ALB actions, while Argo CD is simultaneously tracking those same objects against Git. Without a correctly scoped ignoreDifferences rule and RespectIgnoreDifferences set to true, selfHeal reverts the active canary routing on every reconcile, so a rollout in progress gets silently undone by the tool that's supposed to be delivering it.
The fix is to scope ignoreDifferences as narrowly as the actual noise requires. Ignoring /spec/replicas on a Deployment managed by an HPA is the correct, narrow application of the rule. Ignoring it everywhere, as a blanket policy, quietly masks manual scale changes that genuinely should be flagged. Stateless application resources are where selfHeal earns its keep. Stateful workloads, shared infrastructure, and anything with a legitimate external controller owner need either a scoped ignoreDifferences rule or explicit exclusion from automated sync.
None of this is free at the controller level. The Application Controller reconciles every application in the fleet, comparing desired to live state on each pass, and thousands of concurrent revert loops driven by misscoped selfHeal compound into a measurable performance problem for the controller itself. The 2026 Argo CD User Survey found scaling and performance rank among the most cited challenges teams report running Argo CD at scale.
Sync-wave ordering across a fleet and RollingSync's failure mode
Sync waves order resources inside a single application. RollingSync is the separate mechanism that orders sync operations across applications in a fleet, and its constraints on automated sync introduce a category of drift that has no counterpart at the single-application level. The behavior catches many operators off guard the first time they see it: RollingSync triggers its own sync operations and forces automated sync off on the Applications it generates. Include syncPolicy.automated in a RollingSync ApplicationSet and Argo CD will disable autosync on the generated Applications anyway, logging a warning in the process. Teams that configure both, expecting them to work together, often don't notice the warning until something downstream behaves strangely.
The consequence runs deeper than a configuration quirk. Applications inside a RollingSync fleet stop self-healing between staged rollouts. Drift accumulating in a cluster that hasn't reached its sync wave yet sits there, undetected and unremediated, until the wave finally arrives.
That trade-off exists for a reason, and it's a good one. Deploy the same Application to 80 clusters through an ApplicationSet with no progressive sync, and a misconfigured chart hits all 80 simultaneously. Progressive sync, canary cluster first, then a wave, then the remainder, caps that blast radius at a single cluster while the rest of the fleet waits for the first wave to report Healthy. Argo CD v3.4 brought Progressive Syncs to general availability, and the feature watches for managed Application resources to become Healthy before it advances to the next stage. A drifted or degraded cluster blocks the rollout instead of letting the bad state propagate further.
That same guarantee is also the feature's failure mode. A cluster stuck OutOfSync for reasons that have nothing to do with the rollout itself halts the entire fleet's progress, not just its own. Operators running progressive sync need monitoring on wave progression as a distinct signal alongside health checks on individual applications.
How Argo CD's own remediation undoes incident response
The strongest objection to fleet-wide selfHeal isn't hypothetical. During a live production incident, Argo CD has detected an SRE's emergency kubectl change as drift and synced the old, broken configuration back from Git, undoing the rescue mid-response. SelfHeal did exactly what it was built to do, at the worst possible moment, because it has no way to distinguish a hotfix worth keeping from an accident worth reverting.
Cluster-Level Pause Reconciliation, introduced in Argo CD 3.4, is the direct answer to that scenario, and it's been called the single most important feature in that release. An operator can pause reconciliation for an entire target cluster through the CLI (argocd cluster pause production-cluster) or the UI, without touching global autosync settings and without editing the owning ApplicationSet. Before v3.4, the only workaround was disabling autosync on the Application or ApplicationSet directly, which for generated Applications meant editing the ApplicationSet itself and left no automated path back to re-enabling it once the incident was over.
Cluster pause solves the mechanism; a human still has to make the judgment call about which changes to keep. Auto-sync still can't tell a legitimate hotfix from an accidental change; a human still has to make that call at the moment reconciliation gets paused. Practitioners have built a three-layer pipeline to close that gap: CI catches problems at pull-request time but has no visibility into kubectl edits made after the fact; Kyverno blocks bad configurations at the admission stage but can't revert changes already applied to existing resources; Argo CD reverts runtime drift but can't stop a bad pull request from merging in the first place.
A related failure mode occurs when Argo CD and Flux run in the same cluster. If both tools claim ownership of the same resource in the same namespace, both controllers try to reconcile the same object on their own schedules, and the drift oscillation that follows is close to guaranteed and hard to diagnose after the fact. The operating rule is simple: never let both tools own a resource in the same namespace.
Performance tuning the ApplicationSet and Application Controller as fleet size grows
Everything described so far becomes a resource-allocation problem once fleet size grows past a modest threshold. The Application Controller's reconciliation capacity is the constraint that determines how much drift goes undetected and for how long, and tuning it stops being optional once a fleet reaches meaningful scale. Sharding the Application Controller is disabled by default, but it becomes necessary for multi-cluster deployments: enabling it with a round-robin algorithm and adding horizontal replicas spreads reconciliation load across multiple controller instances instead of concentrating it on one.
The Repo Server carries its own set of tuning levers, and each one addresses a distinct failure. ARGOCD_REPO_SERVER_PARALLELISM_LIMIT defaults to 0, meaning unlimited concurrency, and setting an actual ceiling prevents out-of-memory failures when manifest generation spikes under concurrent load. ARGOCD_EXEC_TIMEOUT defaults to 90s, and large Helm or Kustomize renders need this raised to 180s or higher or the repo server times out and reconciliation stalls.
timeout.reconciliation governs the ApplicationSet controller's own polling interval, and with reliable webhook coverage in place, that interval can be pushed out substantially, to tens of minutes or longer, cutting background load without losing responsiveness where it counts. The right value depends entirely on how much of the fleet actually has webhook coverage versus how much still relies on polling. Repositories carrying thousands of commits also cause Argo CD to sync full history by default, and shallow cloning with depth: "1" cuts repo-server fetch time substantially, at the cost of losing full commit context in the working tree Argo CD checks out. None of these settings fix drift on their own. They determine how quickly the controllers doing the actual detection can keep pace with a fleet that keeps growing underneath them.
Sources
- What Is Argo CD? How GitOps Sync and Drift Detection Work
- GitOps with Argo CD: The Reconciliation Loop That Survives 3 a.m. · Danilo Falcão da Silva
- Building a GitOps Drift Detection & Auto-Remediation Pipeline with ArgoCD, GitHub Actions, and Kyverno | by Coding_Karma | Medium
- Wrong detect drift based on resources field order AFTER applying · Issue #9759 · argoproj/argo-cd
- GitOps with ArgoCD 2026: Cluster Pause, PreDelete Hooks, and the Future of Kubernetes Deployments - DEV Community
- fix(core): ArgoCD self-heal fights KEDA over safezone-worker replicas · Issue #15 · SafeChord/SafeZone-Deploy
- Progressive Syncs - Argo CD - Declarative GitOps CD for Kubernetes
- ArgoCD 3.4 is Coming: 5 Features You Need to Know | by Ramesh | The Platform Engineer | Medium


