What actually happens when identity management breaks down
I used to work at a company where we had a full identity platform handling roughly 40,000 service accounts and about 12,000 human users across four cloud environments. We thought we had it under control until we started seeing the Identity Crisis Identity Crisis problem manifest itself — not as a single outage but as a slow, compounding failure that took us six months to fully diagnose. The issue wasn't that the system failed. The issue was that it succeeded at the wrong things. Here is what I learned dealing with that situation and how to avoid it in your own infrastructure.
Understanding the Identity Crisis Identity Crisis in practice
Most teams encounter this when their identity provider starts producing contradictions between different trust signals. A service account might have valid credentials according to the IAM layer but those credentials have been rotated out of the vault. A human user might show as active in the HR system but disabled in Okta. These mismatches are the core of what people are calling the Identity Crisis Identity Crisis — it is not one problem, it is a class of problems where identity data disagrees with itself across systems. The first thing to understand is that this is almost never a tooling problem. Okta, Azure AD, AWS IAM, GCP IAM, CyberArk — they all work correctly within their own domains. The breakdown happens at the seams between them. I have seen teams spend weeks investigating why their single sign-on was failing, only to discover the real issue was a deprecated SCIM provisioning pipeline that had been silently dropping group assignments for three months.
How I diagnosed the problem in my environment
We started noticing anomalous access patterns. Service accounts were hitting APIs they should not have had permission for. Human users were getting locked out of applications they used daily. The initial assumption was a misconfigured policy. It was not. My approach was to run a full identity reconciliation audit. I wrote a Python script that pulled all principal objects from every identity source — Active Directory, Okta, GCP IAM, AWS IAM, our CI/CD tooling, Kubernetes RBAC, the HashiCorp Vault. That gave me a baseline inventory. Then I cross-referenced permissions at the resource level across each platform. Where a principal existed in one system but not another, or where permissions diverged, I flagged it. The results were worse than expected. Approximately 18% of our service accounts had permission drift. We had orphaned accounts from developers who had left the company. Several automation roles had accumulated excessive privileges through a series of well-intentioned but untracked permission grants during a migration period.
Get the Full Details

The practical fix that actually worked
Once we had the inventory, the fix came down to establishing a single source of truth and enforcing it. We designated Azure AD as the authoritative directory for human identities and GCP IAM as the authoritative source for service accounts in our GCP environment. Every other system was treated as a consumer of that truth, not an independent source. This meant rebuilding our provisioning pipelines. Instead of allowing each team to manage their own access through console clicks — which is how most drift happens — we moved everything through Infrastructure as Code. Terraform modules defined who had access to what. Pull requests became the gate for permission changes. We removed direct console access for production environments entirely. The reconciliation script became a scheduled job that ran every six hours. It would flag any deviation from the declared state and create tickets in our incident tracking system. Most drift was caught within hours now instead of months. The initial setup took about three weeks of focused work. After that, the ongoing maintenance dropped from roughly 20 hours per week of manual auditing to maybe two hours of reviewing automated alerts.
Where this approach falls apart
I want to be honest about the limitations because I have seen people recommend this as a silver bullet and it is not. The biggest issue is that this model assumes you can define a single source of truth. In practice, some environments resist that. Legacy on-premise systems with their own directory services often cannot be easily integrated into a modern SAML/OIDC flow. Third-party SaaS applications sometimes require separate admin consoles with no API access for bulk operations. I had one critical internal tool that only supported basic auth and had no SCIM endpoint, which meant it stayed in the drift zone no matter what we did. Another problem is organizational friction. When you remove the ability for engineers to grant themselves access through a console, there is pushback. Some will call it a security improvement. Others will call it a productivity blocker. The reality is somewhere in between. It does slow down access requests, but it also dramatically reduces the chance of privilege escalation through accidental grants. The tradeoff is usually worth it, but do not pretend it will be popular.
There is also the cost factor. A proper identity reconciliation and IAC-based provisioning setup with tools like Terraform, Atlantis, and a dedicated IAM management layer can run anywhere from $8,000 to $25,000 per month at scale depending on your environment size. For smaller teams with fewer than 500 principals, the overhead may not justify the investment. In those cases, a quarterly manual audit with a simple spreadsheet comparison between your IdP and your cloud IAM consoles is often sufficient.

Common mistakes I see people make
The most frequent error is trying to fix identity drift by adding more tools rather than reducing complexity. Teams will stack on a PAM solution, then add a RBAC layer, then bring in an access review tool, and end up with four systems that all disagree about who has what permissions. The answer is usually the opposite — remove unnecessary identity layers, not add more. Another mistake is treating identity as an IT problem rather than a product problem. The engineers building the services are the ones creating the most dangerous permission drift because they are granted broad access to move fast. I have seen more security incidents caused by developer service accounts with wildcard permissions than by any external attack vector. The fix is not stricter enforcement alone. It is designing the system so that narrow permissions are the default and broad permissions require explicit, documented justification. A third mistake is assuming that SSO solves the identity crisis. Single sign-on is about authentication, not authorization. Just because a user can log in through one portal does not mean their permissions are correct across every application they can reach. We had a case where a terminated employee retained access to a production database through a legacy direct connection that bypassed SSO entirely. The SSO logs showed nothing because the SSO system was not involved in that access path at all.
What I would do differently next time
If I were starting over, I would implement the reconciliation pipeline before building out any new services rather than after. It is much easier to establish clean identity practices from the beginning than to retrofit them onto an existing chaotic environment. We spent the first four months just trying to get visibility into what we actually had. That time could have been spent building controls. I would also have involved the engineering leads earlier in the design process. Our initial rollout caused enough disruption that several teams threatened to stop using the approved provisioning pipeline and go back to manual console access. Once we sat down with them and incorporated their feedback — particularly around making emergency access requests faster — compliance improved significantly. The identity crisis is not going away. As environments grow more distributed and hybrid architectures become standard, the number of identity touchpoints increases faster than most teams can manage them. The teams that handle this well are the ones that treat identity as a foundational concern rather than an afterthought. They accept that it requires ongoing maintenance and budget, and they plan accordingly instead of hoping the problem goes away on its own.