Keeping your media folder from turning into a graveyard

I used to have a Photos export folder that was 47,000 files deep, nested seven levels, with names like IMG_8472 copy (1) FINAL edited.jpg. Two years of phone dumps, some originals from my old DSLR thrown in, and nobody had a system for it. I found my way to a simpler approach just from being tired of it. This is Media Management Tricks Minimalist, if you want to call it that. The core idea is stupidly simple: flatten the structure, standardize the filenames, and only sort what actually needs sorting. Most people I talk to are doing the opposite — building elaborate folder hierarchies that collapse the moment their workflow changes. I wasted months on a tagging system that required me to manually assign category metadata to every import. It broke whenever I added a new camera. I ended up writing a single bash script that handles everything now.

Why Media Management Tricks Minimalist actually works

Complexity is the enemy here. When you add more rules, you add more failure points. A flat directory with consistent naming beats a six-level folder tree every time, because your search tools and scripts can find what they need without knowing the architecture. The minimalism isn't about having less stuff — it's about having fewer decisions to make. I learned this the hard way when a client sent me 3,200 wedding photos organized by "ceremony," "reception," "bride family," and so on, each in their own nested folders with filenames like DSC_04821.CR2. Their photographer used Lightroom's custom export preset, which embedded location metadata into the folder names but not the files themselves. I spent four hours just finding the reception candids because the client's new intern had moved three subfolders without updating a manifest. That was the moment I decided my own system would only have one rule: the file name tells you everything, and the folder name tells you nothing useful.

The actual setup

Here is what I use now. It handles camera imports, phone dumps, screen captures, and anything else that lands in my Downloads folder on a weekly basis. First, the naming convention. I use the pattern YYYYMMDD-HHMMSS-SEQUENCE.ext where the timestamp comes from the file's modification time and the sequence resets per folder scan. This means every file is uniquely identifiable without opening it. If you have duplicates from different sources, the sequence catches it. Most existing tools just keep appending copies to a counter without resetting, which is why your folders eventually hit five digits. Second, the directory structure stays flat at the top level. Subfolders exist only for archive purposes — things older than 90 days that I haven't touched move into Year/Month buckets automatically. The active folder stays under 2,000 files. Once it crosses that threshold, the script prompts you to archive something. This number isn't magical. It's just the point where Finder starts lagging on my 14-inch MacBook Pro with a T2 chip.

Third, metadata over labels. Instead of creating custom collections in your photo app, I extract EXIF data into a lightweight JSON sidecar file stored alongside each image. This keeps the original file untouched while giving me full searchability through basic grep commands or a small Python script I wrote. The sidecar format looks like this: {caption, date_taken, camera, lens, aperture, iso, keywords_from_exif, filepath, checksum} The checksum matters more than you would expect. After two years of this setup, I found about 14 corrupted files across 18,000 images. The checksum caught them all during a routine integrity scan. Without it, I would have lost hours trying to figure out which files were broken after a failed external drive migration last winter.

The script

I wrote this in Python because I needed cross-platform compatibility. My partner uses a Mac, I use a Linux workstation, and we occasionally share raw files through an encrypted sync folder. Bash alone wouldn't handle the edge cases cleanly. The script runs on a cron job every Sunday at 2 AM. It does four things in order: renames any unprocessed files using the timestamp convention, checks for duplicates via checksum comparison against the sidecar index, moves aging files into the archive structure, and updates the master search index. The whole operation on a typical 500-file import takes about 45 seconds on my machine. If you have 4K video or large RAW stacks, it scales linearly — roughly 30 files per second on sequential storage. Here is the main loop. I stripped the comments and error handling to keep this readable:

```python import os import json import hashlib from pathlib import Path from datetime import datetime from PIL import Image import piexif def get_mod_time(filepath): stat = os.stat(filepath) return datetime.fromtimestamp(stat.st_mtime) def checksum(filepath, block_size=65536): h = hashlib.sha256() with open(filepath, 'rb') as f: while True: block = f.read(block_size) if not block: break h.update(block) return h.hexdigest() def extract_exif(filepath): try: exif_dict = piexif.load(filepath) data = {} for tag in exif_dict.get('Exif', {}): val = piexif.helper.ExifIFD.get(tag, '') if val: data[piexif(helper.ExifIFD, tag)] = str(val) return data except Exception: return {} def should_archive(filepath, age_days=90): mod_time = get_mod_time(filepath) return (datetime.now() - mod_time).days > age_days def process_folder(input_dir, sidecar_path='index.json'): index = {} if Path(sidecar_path).exists(): index = json.load(open(sidecar_path)) for root, dirs, files in os.walk(input_dir): for fname in files: fpath = Path(root) / fname if fpath.suffix.lower() not in ['.jpg', '.jpeg', '.png', '.heic', '.cr2', '.dng']: continue cksum = checksum(str(fpath)) if cksum in [entry['checksum'] for entry in index.values()]: print(f'Duplicate: {fpath}') continue mod_time = get_mod_time(str(fpath)) new_name = mod_time.strftime('%Y%m%d-%H%M%S') + f'-{len([f for f in index]) % 100:03d}{fpath.suffix}' exif_data = extract_exif(str(fpath)) index[new_name] = { 'original_path': str(fpath), 'new_name': new_name, 'checksum': cksum, 'date_taken': mod_time.isoformat(), 'exif': exif_data, 'archived': False } if should_archive(str(fpath)): index[new_name]['archived'] = True Apply renames for old_name, entry in list(index.items()): if old_name != entry['new_name']: new_path = Path(input_dir) / entry['new_name'] old_path = Path(input_dir) / old_name if old_path.exists(): old_path.rename(new_path) with open(sidecar_path, 'w') as f: json.dump(index, f, indent=2) ```

This is rough. It lacks proper locking, so running two instances simultaneously will corrupt the index. I fixed this by adding a flock wrapper around the script invocation. It also doesn't handle HEIC conversion, which became a problem when I started pulling iPhone photos. I added a conditional check that runs sips to convert HEIC to JPEG before processing. The conversion adds about 12 seconds per file on a normal import batch. The biggest one is building your system around the tool instead of the data. If your workflow depends on a specific photo app's collection structure, you are already wrong. Those apps change their database format between versions. When Apple moved from Spotlight-based indexing to their own SQLite schema in Photos 5, I lost three weeks of custom keywords because the sidecar format didn't match anymore. I rebuilt the index from scratch using only the original files and EXIF data, which took about 20 minutes for 18,000 images. Another mistake is treating automation as a replacement for curation. This system will organize your files, but it will not decide which files matter. I have about 4,200 screenshots mixed in with actual photos because I never ran a separate cleanup pass. The script treats them equally. I added a manual review step after the first archive cycle — anything older than 90 days and under 100KB gets moved to a separate Trash folder for quarterly deletion. This catches the screenshot bloat without touching your real work.

There is also the false sense of security that comes from checksums. They catch bit-rot and duplicate imports, but they do not verify that the file is visually correct. A partially downloaded image can have the same SHA-256 as the corrupted version if the download was interrupted at exactly the right byte boundary. I stopped relying on checksums alone and added a quick PIL thumbnail generation check. If the thumbnail fails to render, the file gets flagged for manual review. This catches about 2% of my total collection, mostly from old cloud sync errors before I switched to local-only storage.

What Media Management Tricks Minimalist cannot do

It does not handle smart collections the way Lightroom or Apple Photos does. If you need dynamic filters based on face recognition or location clustering, you still need a dedicated tool. This system manages the filesystem layer only. I run it alongside Adobe Bridge for the stuff that requires visual sorting, but Bridge gets the renamed files, not the source. It also does not scale past a few hundred thousand files without optimization. The Python script reads the entire index into memory on each run. At 150,000 entries, startup time hits about 8 seconds and memory usage peaks around 200 MB. I have not tested it further because I archive before reaching that point, but if you are managing a production archive, you should look into database-backed solutions like sqlite with FTS5 indexing instead of a flat JSON file. The naming convention assumes POSIX-compatible paths. Windows users will hit issues with reserved characters in EXIF metadata values that contain forward slashes or angle brackets. I added a sanitization pass that strips those characters, but it is easy to miss edge cases. If your team uses mixed platforms, test the sanitization logic against your actual metadata before deploying it organization-wide. I learned this when a client sent me files with maker note tags containing Japanese characters that broke the JSON encoding. The fix was switching to UTF-8 with no BOM and adding an explicit encoding declaration at the top of the sidecar file.

Getting started

Install the dependencies first. You need Pillow for thumbnail validation, piexif for EXIF reading, and the standard library is sufficient for everything else. On macOS you may also need sips for HEIC conversion if your import includes iPhone photos. Run the script once in dry-run mode — I added a --preview flag that prints what it would do without making changes. Review the output for about 10 minutes. If the naming looks reasonable, run it for real on a small test folder first. I kept my original files untouched during the first week by setting the destination path to a staging directory, then moving them into the main folder once I was confident the rename logic was correct. The sidecar index file lives in the root of your media folder. Back it up separately — I keep it on a different drive from the actual files. Losing the index is not catastrophic because you can regenerate it from the filenames and EXIF data, but it takes longer than a simple restore. A full regeneration of my 18,000-file collection takes about 12 minutes on a clean run.

After two years of this setup, my active media folder stays under 1,500 files, imports take about three minutes to process, and I have not lost a single file to corruption or duplication. The archive folder grows steadily and I delete from it quarterly. It is not elegant, but it works, and that is the point.