Building an Elementary Learners Dictionary That Doesn't Suck
I spent about three years compiling a small corpus-based Elementary Learners Dictionary for ESL beginners. The rough draft ended up being around 4,500 headwords. What I learned mostly has to do with what you leave out rather than what you put in. The first problem people hit is entry selection. You want the words that actual elementary learners encounter, not the words your gut says they should know. I pulled frequency data from the Cambridge English Corpus and the COCA elementary bracket, which is the Corpus of Contemporary American English filtered to roughly A1-A2 output. That gave me a list that looked nothing like what most people assume. Words like "actually," "probably," and "seems" showed up way higher than "beautiful," "delicious," or "however." If you're building a wordlist from your own intuition, you're going to include a lot of decorative vocabulary that beginners will never need and skip over utility words they use every day.
Elementary Learners Dictionary: How to Structure It Differently
Most learner dictionaries follow the standard model: headword, pronunciation, part of speech, definition, example sentence, maybe a collocation note. For elementary level, that standard model creates two problems. Definitions tend to circle back on themselves, and the example sentences often contain words that are above the target level. I solved the definition problem by switching to a controlled defining vocabulary. My entire dictionary was defined using a restricted set of roughly 1,200 headwords. That means every single definition, example, and usage note only used words from that core list. When you do this properly, you catch circular definitions immediately because you can't define a word using another word that isn't in your controlled set. It also means a true beginner can read every definition without needing a separate reference work. The example sentence problem is harder. I fixed it by running every example through a simple frequency filter. If an example sentence contained more than two words outside the elementary bracket, I rewrote it. This sounds straightforward until you hit phrasal verbs, where the particle itself might be low-frequency but is completely essential to the meaning. With "look up" as a phrasal verb, writing "Please look up the word in the dictionary" uses only elementary vocabulary but teaches the wrong sense. Writing "She looked up her friend when she arrived at the station" teaches the right sense but introduces "arrived" and "station," which aren't elementary words. I ended up adding a separate phrasal verb section where the examples were allowed slightly more complex vocabulary because the phrasal verb itself was the learning target. The tradeoff was worth it.
Practical Entry Design Decisions
Headword selection is where most people waste the most time. I initially tried to include every useful verb, which ballooned the project. Then I narrowed to a core list of roughly 600 verbs based on corpus frequency at the A1-A2 range. Nouns came to about 1,800. Adjectives around 900. The rest — adverbs, prepositions, conjunctions, pronouns — was mostly fixed and took up about 500 entries. The total landed somewhere around 4,000 to 4,500, which is a manageable size for a digital product or a print reference book at this level. One thing that surprised me: the most contentious entries were things adults take for granted. Words like "some," "any," "enough," and "very" are extremely high frequency at the elementary level, but they are also extremely hard to define clearly using a controlled vocabulary. "Very" seems simple until you try to define it without using other intensity adverbs. I ended up using a functional definition: "very is used to say that something has a quality to a large degree." It's not elegant, but it's accurate and it uses only elementary words. I did similar pragmatic definitions for "some" and "any" based on countable versus uncountable noun contexts rather than trying to give abstract semantic definitions.Get the Full Details

Collocations and Chunks at Elementary Level
Another decision that matters more than people think: whether to include collocations. Traditional learner dictionaries treat collocations as intermediate-level material. But corpus data shows that elementary learners produce a lot of unidiomatic phrases because they're translating word-for-word from their L1. "Make a photo" instead of "take a photo" is the classic example, but there are dozens of others that come up constantly in classroom settings. I included a small collocation section for the top 400 verbs only. Each verb entry listed its three most frequent collocations, drawn directly from the corpus. This meant "make" got "make a decision," "make a mistake," and "make money," while "do" got "do homework," "do business," and "do a favor." The whole collocation apparatus added maybe 800 lines to the dictionary. The payoff was significant — students stopped making the most common error patterns, and teachers reported fewer corrections on the same mistakes.A Specific Problem I Ran Into
Here is a concrete edge case that nearly derailed the project. I was finalizing the entry for "book" as both a noun and a verb. The noun is trivial. The verb, though, has multiple senses: to reserve something, to record something in writing, and the older sense of putting someone in prison. In a corpus search, the verb "book" appeared frequently in British English contexts meaning "to arrest" — "He was booked for vandalism." That sense doesn't belong in an elementary dictionary. But filtering it out required me to manually review every occurrence in the corpus data, which for "book" alone was over 200 entries across multiple years of text. I built a simple script that flagged verb usages where the surrounding context contained words from a pre-built list of advanced vocabulary (like "arrest," "charge," "suspect," "police" in legal contexts). It caught about 90 percent of the non-elementary verb uses. The remaining 10 percent I reviewed by hand. The whole process took about six hours for one headword. If you're building a dictionary like this solo, budget that kind of time for high-frequency polysemous words. They will eat your schedule.Common Pitfalls and Where This Approach Fails
There are genuine limitations to building an Elementary Learners Dictionary from scratch this way. The controlled defining vocabulary approach works well until learners need to discuss abstract topics — politics, science, technology — where even elementary-level vocabulary isn't sufficient. A student reading a simplified news article might encounter "government," "election," or "policy," and your dictionary won't help them because those words fall outside your core list. You can expand the defining vocabulary, but then you lose the pure beginner readability that made the approach worthwhile in the first place. Another limitation: pronunciation guides. If you're targeting a global audience, you need at least two phonemic transcriptions — one for British English and one for American English. Doing this consistently across 4,500 entries by hand is tedious and error-prone. I used a TTS-based pronunciation generator to create the initial transcriptions, then manually corrected entries where the automated system produced inconsistent IPA symbols or misidentified stress patterns. About 15 percent of the entries needed manual correction. If you're not comfortable with phonetics yourself, this is the part where you should either learn the basics or find a native-speaking phonetics consultant.
And here is the blunt truth about the end product: an Elementary Learners Dictionary is inherently incomplete. No matter how carefully you build it, there will always be words that advanced beginners know and elementary beginners don't, or vice versa. The line between elementary and pre-intermediate is fuzzy, and any frequency-based cutoff you draw will feel arbitrary to someone on either side. The best you can do is be transparent about your methodology and update the wordlist regularly as new corpus data becomes available. If you're considering building your own, I'd recommend starting with a smaller corpus — maybe 500 to 1,000 entries — and testing it with actual elementary learners before expanding. Classroom feedback will surface issues that no amount of corpus analysis catches, like entries that look useful on paper but never actually appear in the kinds of texts your target students read.