Extracting Social Networks from Text with Our Mutual Friend
Most people who stumble onto this library are coming from a digital humanities background or are working on some kind of character relationship mapping project for fiction. The tool sits somewhere between simple keyword extraction and full NLP pipeline, and honestly, that's both its strength and its weakness. I've used it on a few projects over the years, mostly for pulling relationship data out of Victorian novels, and I'm going to walk through what it actually does, how to get it running, and where it tends to trip people up. The core idea is straightforward. You feed it a text, and it identifies named entities and then determines the relationships between them based on co-occurrence patterns, grammatical structures, and configurable relationship types. Think of it as turning a novel into a graph dataset you can actually query. Under the hood, it uses dependency parsing from spaCy and runs its own co-reference resolution to figure out when "he" and "Mr. Darcy" are the same person in a given passage. The output is typically JSON or CSV representations of nodes and edges, which you can then pipe into Gephi, NetworkX, or whatever visualization tool you prefer. I usually go straight to NetworkX for the initial analysis and export to Gephi only when I need publication-quality graphics.
Installation is simple enough. You can grab it from PyPI if your project uses Python 3.8 or later: pip install our-mutual-friend I'd also recommend installing the spacy English model separately, either en_core_web_sm for speed or en_core_web_trf if you're processing longer documents and need better accuracy on tricky syntax. The difference matters more than you might expect on dense prose.
Basic Usage
Here's what a minimal script looks like in practice: from our_mutual_friend import extract_relationships\n\nresult = extract_relationships(\n text=document_text,\n relationship_types=["family", "romantic", "professional", "hostile"],\n min_co_occurrences=3\n)\n\nprint(result.to_json()) The relationship_types parameter is where most beginners waste time. By default, the library scans for a fairly broad set of relation categories. If you're working with a specific genre like Victorian novels, you'll want to narrow this down. I spent about three weeks debugging weird false positives before I realized my relationship type configuration was pulling in coincidental co-occurrences as "professional" ties between characters who never actually interacted in any meaningful way.
Get the Full Details

The min_co_occurrences threshold is your primary control for noise. Set it too low and you get a graph that looks impressive but is mostly garbage. Set it too high and you lose genuine but sparse relationships. For most literary texts, I find 3 to 5 is the sweet spot depending on the length and density of the source material.
Edge Cases and What Went Wrong For Me
Here's the thing nobody mentions in the README. When you're processing texts with a lot of dialogue, the co-reference resolution starts breaking down in predictable ways. I ran into this specifically with a corpus of serialized fiction where characters were frequently referred to by title alone ("the gentleman," "the young lady") without clear antecedents in nearby sentences. The parser would arbitrarily assign these pronoun-like references to whichever entity had the highest node weight at that point in the document, which created entirely spurious edges in my network. The workaround I ended up using was a two-pass approach. First pass: run the standard extraction with a higher min_co_occurrences threshold to establish baseline relationships. Second pass: filter out any edges that involved entities referenced only through ambiguous third-person descriptors by cross-referencing with a manual entity list. It added maybe an hour of preprocessing work but eliminated roughly 40% of the false positive edges I was seeing in the output graphs. Another issue I hit repeatedly: the library struggles with texts that have significant narrative frame shifts. If a novel has a prologue set decades before the main story, the parser treats all characters as potentially co-existing across the entire timeline. I learned this the hard way when my relationship graph showed romantic ties between a grandmother character and a grandson character, which was biologically and narratively impossible but technically accurate according to the co-occurrence data.
The fix there was chunking the document by section and running separate extractions, then merging the results with a temporal validity filter. Not ideal, but it saved the project.

Performance Considerations
This isn't a fast library. Processing a standard 80,000-word novel with the transformer-based spaCy model takes roughly 8 to 12 minutes on a modern laptop. The small model version cuts that to about 2 to 3 minutes but sacrifices some relationship accuracy, particularly on ambiguous sentences. For corpora of multiple novels, I'd suggest setting up a queuing system rather than running everything sequentially. I use Celery for this, and it typically handles batch processing without major issues. Memory usage scales roughly linearly with document length. I've processed documents up to about 200,000 words before hitting OOM errors on an 8GB machine. If you're working with very long texts, split them into chunks of around 50,000 words and process separately.
When This Approach Fails Completely
There are texts where Our Mutual Friend simply will not give you reliable results, and it's worth knowing these upfront so you don't waste weeks on a dead end. Highly experimental or postmodern fiction with unreliable narrators, shifting perspectives, and intentional ambiguity in character identification produces mostly noise. The parser assumes a stable referential system, and when that breaks down, the output becomes statistically indistinguishable from random. Texts with extremely large casts also present problems. I tried running it on War and Peace and ended up with a graph so densely connected that individual relationship edges lost all interpretive value. The library doesn't have built-in community detection or edge-weight normalization that would help here, so you're essentially on your own for post-processing. If you're working with these kinds of texts, I'd recommend looking at alternatives like Stanford's CoreNLP pipeline with custom relationship extraction models, or if you're in an academic context, the open-source pyrelate package which has some more sophisticated handling for ambiguous reference resolution. Neither is a drop-in replacement, and both have their own learning curves, but they handle edge cases that Our Mutual Friend just glosses over.
The library itself is maintained on GitHub under the name our-mutual-friend, and the documentation covers the API surface adequately but doesn't address most of the practical issues I've described. That gap between the documentation and actual usage is why I'm writing this. If you're starting a project with it, budget extra time for cleaning up the output rather than assuming the raw results are analysis-ready.
