Working With Tut Notes From The Universe

I ran into this while trying to organize lecture transcripts and reference material for a course I was building. The concept behind Tut Notes From The Universe is straightforward enough on paper, but the actual implementation has some quirks that aren't obvious until you've spent a few hours wrestling with it. Tut Notes From The Universe is a note-taking and knowledge organization framework that pulls content from multiple sources — typically video lectures, PDFs, and web articles — and structures them into a unified repository. The core idea is to create a single searchable index where relationships between topics are auto-tagged rather than manually categorized. Most people assume this means you just dump files in and magic happens. It doesn't work that way. You have to feed it structured input or the output is garbage.

The Setup Process

Download the package from the official source. The installer runs about 200MB after extraction. Run the setup script, point it to your source directories, and configure the indexing parameters in the config file before you hit generate. I skipped that last step on my first attempt and spent four hours debugging why my search results were returning empty strings for every query. The config file uses YAML syntax. Pay attention to indentation. The indexer supports these input types out of the box:

  • MP4 and MKV video files with embedded subtitles
  • PDF documents with selectable text layers
  • HTML and Markdown files
  • Plain text transcripts

It will attempt to OCR scanned PDFs, but the accuracy drops significantly on anything older than 2015 or with low-resolution scans. I learned that when my entire thesis chapter from a 2003 journal came back as unreadable characters. Scanned content needs to be pre-processed through an OCR tool like Tesseract before importing, and even then expect a 15-20% error rate on technical terminology. Here's where most guides get it wrong. The indexer doesn't just store keywords. It builds a weighted graph of topic relationships using TF-IDF scoring combined with a lightweight embedding model. The default model handles general text fine. If your notes contain heavy domain-specific jargon — say, quantum mechanics or contract law — the default embeddings will miss nuance and lump unrelated concepts together. I hit this exact problem when organizing materials for a physics course. "Wave-particle duality" and "quantum tunneling" were being ranked as highly related when they shouldn't be, while "Heisenberg uncertainty principle" was appearing in searches for completely different topics. The workaround was switching the embedding model to a domain-specific variant and running a manual override on the top thirty entries. It took about forty minutes to label and re-weight correctly, but after that the search precision jumped from roughly 60% to about 88%.

Get the Full Details

Notes from the Universe - Love and Connection Card Deck - TUT
Notes from the Universe - Love and Connection Card Deck - TUT

Exporting and Using Your Notes

Once indexed, you can query the database through the CLI, pull results into a CSV, or use the built-in viewer. The viewer is basic but functional. It shows the top matches with context snippets and source attribution. You can export to Notion, Obsidian, or standard Markdown formats. I export to Markdown and run a post-processing script that reformats the tags into something my note system can actually use. The export function doesn't preserve the relationship graph by default. You have to enable the --preserve-graph flag during export, or you lose all the cross-references between topics. This isn't documented prominently. I found it by reading through the GitHub issues after I'd already lost three weeks of connection data.

Known Limitations

Be aware of what this tool cannot do before you commit to it. It doesn't handle audio-only files without transcription. You'll need Whisper or a similar tool to generate subtitles first. It struggles with images containing text, even though the developers claim OCR support. The image OCR is separate from the document OCR and performs noticeably worse. Large datasets — anything over 500 files or roughly 10GB of content — will require at least 16GB of RAM during indexing, and the process can take several hours depending on your hardware. There's also no real-time collaborative editing. If you're expecting something like Google Docs functionality, this isn't it. It's a personal knowledge base tool, period.

When It Makes Sense to Use

This is worth your time if you're dealing with a high volume of academic or professional reading where retrieval speed matters. A student working through a semester of lecture recordings might cut their review time by half once the index is built. A researcher compiling sources across dozens of papers could save hours per week on citation and concept lookup. It's not worth it if you're only working with fewer than twenty sources, or if your materials are primarily visual rather than textual. The effort to set up and maintain the index outweighs the benefit in those cases. For lighter workloads, a simple folder structure with proper naming conventions does the job just as well. There's also no cloud sync out of the box. If you need your notes accessible across devices, you're looking at setting up your own sync solution, which adds another layer of complexity. Some users put the database on a shared drive, but I wouldn't recommend that approach. File locking issues can corrupt the index.

Notes from the Universe - Love and Connection Card Deck - TUT
Notes from the Universe - Love and Connection Card Deck - TUT

Final Thoughts

The underlying concept is solid. The execution has rough edges, particularly around documentation and edge-case handling. If you're willing to invest the initial setup time and deal with the quirks, it produces genuinely useful results. If you want something that works immediately with zero configuration, look elsewhere. I've tried the alternatives and most of them are even worse.