International Id Checking Guide
Identity verification across borders is one of those things that sounds simple until you actually have to do it. You have applicants from twelve different countries, each with their own document formats, languages, and validation rules. The easy part is uploading a passport or national ID. The hard part is knowing what a valid Mexican INE looks like versus a forged one, or whether a Brazilian RG is still issued in most states or if everyone switched to CNH by now. I ran a verification pipeline for a fintech client last year handling transactions across Latin America, Southeast Asia, and Eastern Europe. We processed roughly 4,000 identities per month. About six months in, I learned that the OCR confidence scores we were trusting were lying to us on Vietnamese CCCD cards. The system reported 94% confidence on a document that turned out to be a photoshopped version someone submitted through our mobile app. The OCR parsed the right fields, right numbers, even the right address. It just didn't catch that the holographic seal had been digitally rendered instead of physically printed. That one cost us about three weeks of debugging before we started cross-referencing document security features through image forensics rather than relying purely on text extraction.
Where to Start With International Id Checking Guide
The first decision you need to make is whether you are building this yourself or buying it. A lot of people jump straight into building because they think they can save money. You will not save money unless you have a team that already understands document forensics, regional ID systems, and the compliance requirements of every country you operate in. The vendors that actually do this well charge between $0.05 and $0.50 per verification depending on volume and complexity. For anything under 500 verifications a month, the vendor route is almost always cheaper when you factor in engineering time. If you are building, you need to understand that "international ID checking" is not one system. It is a collection of subsystems, each tuned for a specific region's documents. A single pipeline cannot handle a Japanese My Number Card, a Nigerian NIN slip, and a German Personalausweis with the same accuracy without separate configuration for each. The common architecture breaks down into these layers: document type detection, liveness or authenticity checks, OCR and data extraction, database or registry cross-reference, and risk scoring. Document type detection is where most people underestimate the difficulty. A Polish dowod osobisty and a Belgian identity card look nearly identical to a basic CNN. You need either a trained model on actual document images or a vendor API that already has this coverage. I spent two months trying to get a custom ResNet-50 classifier to reliably distinguish between an Estonian ID-card and a Lithuanian one. The color schemes are too similar and the layouts nearly mirror each other. We ended up routing those through a vendor API and kept the custom model only for documents where the vendor had poor coverage, like obscure provincial IDs from Central Asian countries.
The Liveness and Authenticity Layer
This is the part that separates people who have been caught red-handed from people who haven't. Uploading a photo of a screen showing someone else's ID is trivially easy. Asking for a selfie with a random gesture doesn't solve that either, because deepfake video can now produce convincing responses to arbitrary prompts in real time. The approach that actually works is combining multiple signals. You want a liveness check that uses challenge-response with unpredictable timing, infrared or near-infrared imaging to detect screen replay, and texture analysis to catch printed photographs held up to a camera. For document authenticity, you check for microtext presence, security thread patterns, UV-reactive element detection through appropriate lighting, and consistency between the MRZ (machine-readable zone) and the visual zone data. If the MRZ says the passport number is different from what appears in the visual field, that is an automatic reject regardless of anything else. One thing nobody tells you about this: the MRZ-to-visual-data consistency check catches about forty percent of fraud attempts on its own. Most people skip it or implement it poorly because it feels too simple. A forged Indonesian KTP will often have manually entered data that contains typos or mismatched formatting compared to the printed fields. The MRZ itself is machine-generated and extremely difficult to fake correctly by hand.
Get the Full Details

Registry and Database Cross-Reference
This is the layer where "international" actually matters. A US passport check can query DHS databases. A German ID check can hit the Zentralregister. But most countries do not have publicly accessible or even API-accessible identity registries. When you verify an Indian Aadhaar card, you are either using the UIDAI's official API with user consent and OTP verification, or you are relying entirely on document appearance analysis, which is significantly less reliable. The countries where registry access is easiest: United States, United Kingdom, Canada, Australia, most EU nations through the eIDAS framework. The countries where you will mostly be blind: India (Aadhaar has strict consent gates), Brazil (CPF validation exists but is limited), Indonesia (no public NIK lookup), Philippines (no centralized ID verification API for private companies), and most of Sub-Saharan Africa where national ID systems are still being digitized. For these regions, you fall back on third-party data brokers and alternative verification methods. Phone number ownership checks, email velocity checks, device fingerprinting, and address verification through postal or utility data. None of these are as strong as a government registry match, but they add enough signal to catch the casual fraudster. A determined fraudster with access to real credentials from a compromised individual will still pass all of these checks.
Pipeline Implementation Notes
I would structure the pipeline to process documents in parallel across these stages rather than sequentially. Document type detection and liveness check can run simultaneously. OCR extraction happens alongside image forensics. Only the final risk scoring needs to wait for all previous layers to complete. This cuts total verification time from around forty-five seconds down to roughly twelve on a well-provisioned cloud setup. Storage and retention are a separate headache you need to plan for upfront. GDPR requires that you not retain biometric data longer than necessary. Some countries mandate that identity documents be deleted within a specific timeframe after verification. Others require you to keep records for seven years for anti-money laundering compliance. These requirements conflict with each other regularly. I had a client who operated in both Germany and the UAE and spent three months writing a policy engine that determined which retention rule applied based on the user's country of residence and the type of data collected. The solution was to store raw document images in an encrypted bucket with automatic deletion schedules tied to jurisdiction, while keeping only the verified data fields and a hash of the original document for audit purposes.
Common Failure Modes
Here is what goes wrong in production. Glare on laminated IDs causes OCR to miss entire lines of text. This is especially bad with Middle Eastern IDs that often have reflective laminate over the photo area. The workaround is to require users to tilt the document slightly or to use multiple capture angles. Some vendors now include a guidance overlay that detects reflections in real time and asks the user to adjust lighting before capturing. Expired documents are accepted as valid because the data itself is internally consistent. An expired Philippine driver's license will parse perfectly fine. You need to explicitly check the expiration date against the current date and treat expired documents as a separate risk category, not a rejection. Some jurisdictions legally require expired documents to be flagged rather than rejected outright because the person may genuinely not have renewed yet. Transliteration errors in names cause matching failures downstream. A Turkish name like "Ğıbrel" gets mangled by any OCR system into "Gıbre1" or "Gilrel" and then fails to match against any database record. The fix is to implement fuzzy matching with character-level edit distance that accounts for diacritical marks and known transliteration patterns for each language. This adds maybe ten percent to your compute cost but prevents a twenty percent false rejection rate on non-Latin script documents.

A Word on Vendor Selection
If you buy rather than build, test every vendor against your actual applicant demographic mix before signing a contract. Most vendor documentation shows accuracy numbers based on a balanced dataset of common documents from developed countries. That is not your dataset. You need to run a proof of concept with at least five hundred real submissions from each region you serve. I learned this the hard way when a vendor we chose reported 99.2% accuracy on passport verification. Their testing set was ninety percent US and UK passports. When we launched with mostly Vietnamese and Pakistani passports, accuracy dropped to seventy-one percent. We switched vendors within two weeks and ate the integration cost. No system can reliably verify identities from countries with no digitized record-keeping infrastructure. If you are trying to verify someone from rural South Sudan or parts of Myanmar where national ID systems do not exist or are not accessible, you are working with incomplete data regardless of how good your technology is. The honest answer in these cases is to fall back on manual review with a human who speaks the local language and understands the local document system, or to decline the application. Automating through uncertainty produces false positives at scale, and false positives in identity verification are worse than false negatives because they let fraudulent identities through rather than blocking them. The International Id Checking Guide is not a single tool or product. It is a set of decisions about which documents to accept, which signals to trust, and which risks to absorb. Get those decisions right first. The technology comes after.