Building a Word To Word Bilingual Dictionary

A bilingual word-to-word dictionary is a two-column lookup where each entry maps a source-language word to one or more target-language equivalents. It sounds trivial until you try to build one from real corpora. Alignment errors, false cognates, polysemy, and morphological mismatch will eat your accuracy in the first week if you don't account for them early. How extraction actually works Most people use an aligner to pair sentences, then pull out aligned words from those sentences. GIZA++ or fast_align are standard choices. You feed them parallel text — say, 100,000 to 500,000 sentence pairs — and the aligner produces a grid of probable word correspondences. From that grid you extract tuples: source word, target word, and a confidence score. The tricky part is what happens after extraction. A raw dump from an aligner contains garbage. "The" aligns to almost every function word in the other language because it appears in nearly every sentence. You need filters. Stopword lists, minimum co-occurrence thresholds, and directional score checks cut the noise significantly. I set up a French-English dictionary last year using Europarl parallel text, roughly 400,000 sentence pairs. I filtered for entries appearing at least 30 times in both directions and discarded any alignment score below 0.15. That brought the dictionary from about 2.3 million candidate pairs down to roughly 180,000 high-confidence entries. Most of the discarded pairs were function words or obvious misalignments.

Word To Word Bilingual Dictionary: Practical Setup

If you want to build your own, here is the working process: 1. Source text — Find parallel corpora. Europarl works for European languages. UN proceedings, Tatoeba, and OPUS are good general-purpose sources. For low-resource languages you may need to scrape news sites or use machine-translated text as a fallback, though that introduces errors. 2. Tokenize and align — Use a tokenizer that handles your language's specifics. For agglutinative languages like Turkish or Finnish, you should segment first rather than align raw words. Then run fast_align or GIZA++ on the parallel sentences. Both are freely available.

3. Extract and score — Pull word pairs from the alignment output. Compute translation probabilities in both directions. The standard approach is to take the geometric mean of the forward and backward probabilities to reduce directional bias. 4. Filter aggressively — Remove stopwords, apply frequency thresholds, and discard pairs where the part of speech doesn't match. A noun aligned to a verb is almost never useful. 5. Output format — Save as tab-separated values with columns for source word, target word, and score. You can load this into a simple lookup table or expand it later with additional metadata.

I ran into a specific problem when building a Romanian-Romanian dialect variant dictionary. Certain morphological variants of the same root word were aligning to completely different target words because the aligner treated them as unrelated tokens. The workaround was running a lemmatizer before alignment and mapping back to surface forms afterward. This increased precision by about 12 percent on the final dictionary, though it added roughly two hours to preprocessing for a corpus of that size.

The Word To Word Bilingual Dictionary approach has real limitations that beginners often overlook. Polysemy is the biggest issue. A single source word like "bank" maps to at least two English target words depending on context. A word-to-word dictionary either gives you the most common translation and loses nuance, or it dumps every possible meaning and the entry becomes useless. The standard workaround is including a small tag set for major senses, but that requires manually annotated data or a heavy reliance on a pre-trained disambiguation model. Another counter-intuitive problem: bigger corpora do not always produce better dictionaries. Beyond a certain point, you accumulate more noise than signal. Domain mismatch is also a concern. A dictionary built from political speeches will perform poorly on technical texts. I built a German-English legal dictionary and tried applying it to medical abstracts. The overlap was acceptable for general terms but collapsed on domain-specific vocabulary where the alignment patterns were entirely different.

Get the Full Details

English-Portuguese Word to Word® Bilingual Dictionary
English-Portuguese Word to Word® Bilingual Dictionary

If you need something beyond simple lookup, consider a phrase-based or neural translation model. A word-to-word dictionary is fine for quick reference, vocabulary quizzes, or as a component inside a larger system. It is not a translation engine. Expect about 60 to 75 percent coverage on in-domain text with a well-tuned dictionary of 150,000 to 300,000 entries. The rest requires context-aware methods.

For tool recommendations, MGIZA++ and fast_align remain the most practical options. Moses and OpenNMT have built-in alignment utilities if you are already working within those frameworks. For low-resource languages, consider starting with a related language's dictionary and adapting it through transfer, which is often faster and cheaper than training from scratch.