What This Actually Is
Identity Recover Guide is a framework I've been building and refining over the past few years for handling orphaned identities in IAM environments. The basic problem it solves is straightforward: when someone leaves a company, their accounts don't just disappear. SaaS tools, internal apps, legacy systems — they all keep the account alive somewhere, and eventually you end up with hundreds of ghost accounts sitting there, consuming seats, triggering security flags, and making audits a nightmare. The guide lays out a systematic approach to identifying these accounts, classifying them, and either recovering the identity data or decommissioning it cleanly. It's not a single piece of software you download and install. It's more like a playbook with scripts and configuration templates you can adapt to your environment.
Identity Recover Guide Setup and Configuration
You'll need a few things before you start. A SCIM-enabled identity provider is the ideal base, but this works with LDAP and Active Directory connectors too. You need read access to your IdP's audit logs and at least read-only access to each downstream application's user directory. I usually set this up on a small Linux VM — Ubuntu Server, 2 vCPUs, 4GB RAM is plenty. The automation scripts are Python-based and run locally, not in the cloud, which keeps sensitive identity data from leaving your network. The first step is connecting to your IdP. If you're on Okta, the API calls are well-documented. For Azure AD, you'll need the Microsoft Graph SDK. Here's what the initial connection script looks like: I typically write a small bootstrap script that pulls all active users with their last login timestamp. From there, I cross-reference against a list of known leavers from HR exports. The mismatch between those two datasets is where the orphaned identities live.
One thing that catches people off guard is the timing. If you try to pull historical login data from your IdP, most providers only go back 6 to 12 months. That's fine for most companies, but if you have contractors who rotate every 18 months, you'll miss a chunk. I solved this by piping the IdP query into a local SQLite database on the first run, then incrementally updating it daily. Now the history is complete and local.
Get the Full Details

How to Run the Recovery Pipeline
Once your connections are solid and your orphaned identity list is built, the pipeline moves through three stages: classification, recovery attempt, and final disposition. In the classification stage, each orphaned account gets tagged based on a set of rules. Accounts that haven't logged in for over 90 days and belong to employees who left more than 6 months ago go into the "decommission" bucket. Accounts that are still active or were recently used get flagged for manual review. Shared service accounts and automated service identities get their own category because they should never be touched without explicit approval. I ran into a specific problem last year that took me about two weeks to figure out. We had a batch of roughly forty accounts tied to a deprecated internal project. All the project leads had moved on, but the accounts were still showing as "active" in the IdP because a background service was using stored credentials to poll an API every twelve hours. Standard inactivity rules didn't catch them. The workaround was to check the API call logs directly instead of relying on the IdP's login timestamp. The service was authenticating with OAuth tokens, and those tokens hadn't expired even though the humans behind the project were gone. I identified the service principal in the token metadata, rotated the credentials, and the accounts finally stopped showing as active.
The recovery stage attempts to automatically disable or terminate accounts in each downstream application. This is where the SCIM provisioning paths matter. If your IdP uses SCIM 2.0 properly, you can send a PATCH request to disable an account and most modern SaaS tools will honor it. For older systems without SCIM support, you need to fall back to API calls or even scripted UI automation, which is slower and more fragile.
Limitations and Where This Falls Apart
This approach works well for modern, cloud-native SaaS environments. It breaks down pretty quickly in a few specific scenarios. If you have on-premise applications that authenticate directly against Active Directory without any intermediary layer, the script can't easily reach them. You'll need to run the cleanup logic inside your AD domain or use a connector tool. Legacy systems that store passwords locally and don't integrate with your IdP at all are completely out of scope. Nothing in this guide will touch those. You'll need a separate asset inventory process for that. Another issue is multi-tenancy. If your organization serves multiple business units or subsidiaries with separate IdP instances, you need to run this pipeline independently for each tenant. The scripts don't aggregate across tenants automatically. I've seen teams try to force a single pipeline to handle all tenants and end up with partial results because one tenant's API throttles the others.

The biggest limitation, honestly, is that this only works if you have visibility into your identity landscape in the first place. If you have shadow IT — departments signing up for tools without going through procurement — no script will find accounts that don't exist in your IdP. The recovery guide can only clean up what it can see. Regular audits of SaaS spend and mandatory procurement workflows are the real fix for that.
Practical Tips From Experience
Don't run the full pipeline on a Monday morning. If something goes wrong — like a misconfigured rule accidentally disabling a critical shared account — you want people available to respond. I schedule these runs for Thursday afternoons so there's time to catch issues before the weekend. Keep a full backup of your identity data before touching anything. Export the user list from your IdP, along with all group memberships and role assignments. I learned this the hard way after a bad export script corrupted a local copy of our staging environment user data. Took me six hours to restore from the backup. The scripts and configuration templates are available through the Identity Recover Guide repository. It's a GitHub repo with the full documentation, Python scripts, and example configurations for Okta, Azure AD, and OneLogin. No license restrictions, just clone and adapt.
If you're dealing with a very small environment — fewer than fifty users across three or four SaaS tools — this might be overkill. The manual process of checking each app and deactivating accounts yourself could take you maybe half a day upfront and then twenty minutes per month for ongoing cleanup. The guide pays for itself once you're past roughly thirty users and ten connected applications.

Identity Recover Guide Common Pitfalls to Avoid
People tend to rush the classification rules. Setting the inactivity threshold too low — say, 30 days — will flag legitimate temporary absences. Parental leave, sabbaticals, long conferences. I set mine at 90 days and add a manual review step for anything between 60 and 90 days. Another mistake is assuming that disabling an account in the IdP automatically disables it everywhere. It doesn't. SCIM propagation has a delay, usually between 5 and 30 minutes depending on the connector, and some applications cache session tokens for much longer. An account might show as disabled in your dashboard but still have an active session on a user's machine. Finally, don't skip the reporting step. After each run, generate a summary of what was deactivated, what was flagged for review, and what failed. I send this to the IT operations team and file it for compliance audits. Having a paper trail matters more than you'd expect when someone asks why an account was deleted three months ago.