Working With Oscar Winner Data Year by Year
I spent three years pulling and cross-referencing Academy Award data for a production company that needed to brief clients on historicalBest Picture trends before awards season. What followed was an unglamorous process of matching incomplete datasets against each other and learning where the public records quietly contradict themselves. The main resource most people end up using is the official Oscars database at oscars.org, but that site does not expose clean structured data. It is built for casual browsing, not for bulk extraction. If you are doing this manually, you will hit rate limits on any scraper and eventually get IP-blocked. If you are hiring someone to scrape it for you, they will send back half the fields you asked for because the site layout changes without warning and some winners are listed under alternate titles depending on the year.
Where to Find Oscar Film Winners By Year
The simplest path depends on what you need the data for. If you want accuracy for a professional deliverable, the Academy itself publishes a downloadable PDF list for each ceremony, usually within a week after the event. Those PDFs are the ground truth and they include the exact film titles as submitted, which matters because the public release title and the submission title are not always identical. If you need a programmatic feed, the IMDb API and the Academy Awards dataset on Kaggle are the two options people actually use. The Kaggle dataset is a cleaned CSV version compiled by a user named nikhaldini, last updated in early 2024. It covers every category from 1929 through 2024 and includes the win versus nomination flag, category, and movie title. The download is free. The catch is that the title normalization is inconsistent, the year column sometimes conflicts with the ceremony year, and the Best Picture field has duplicate entries for co-winners like Parasite that look like two separate rows but are really one winner. I ran into this exact problem when a client asked for a clean pivot table of every Best Picture winner with its release year, budget, and domestic gross. The dataset had the release year listed as the award year for several films because the Academy counts eligibility differently than the MPAA distribution date. Green Book, for example, shows a ceremony year of 2019 but a wide release year of 2018, and the dataset labels it ambiguously. My workaround was to write a small Python script that matched each winner against the Academy's official eligibility year list, then cross-checked against Box Office Mojo for the actual theatrical release. That reduced the mismatch error rate from about fourteen percent down to roughly one percent.
The Category Trap Most Beginners Miss
The Oscars have far more than Best Picture, and anyone trying to build a reliable dataset for "Film Winners By Year" usually underestimates how much the category roster has shifted over time. Short Film, Documentary Feature, Animated Feature, International Feature, and several craft categories appeared at different points in history. Animated Feature did not exist until 2001. Before that, animated films were not eligible for Best Picture unless they also qualified through live-action pathways, which is why no Disney film won before Beauty and the Beast was nominated in 1992 despite being released in 1991. Another issue people encounter is the tiebreaker rule. Ties have happened in Best Picture multiple times. My Darling Clementine did not win; Going My Way won in 1944. An Officer and a Gentleman beat out four other nominees in 1982. The most confusing tie was 2013, when Argo won over Django Unchained, Les Misérables, and Silver Linings Playbook—no tie there, just a crowded field. The real tie confusion comes from categories like Best Actor, where Jackson and Pacino tied in 1993, and Brando and Hoffman effectively shared honors in earlier years before the formal tie rule was codified. For film-level data, the 1928/29 ceremony is its own edge case because it honored two separate films: Sunrise for Unique and Artistic Picture and Wings for Outstanding Picture, with the latter being the retroactive Best Picture honor. When people ask for Oscar winners by year starting in 1928, they almost always mean Wings, not both.
Get the Full Details

How to Actually Build a Reliable Yearly Dataset
Start with the Kaggle CSV, then layer the official Academy PDF for the most recent ceremonies where the CSV may lag. I keep a local SQLite database and join the two sources on film title plus year, using a fuzzy match column for titles that differ slightly between sources. For the Best Picture column specifically, I normalize the title against the Academy's submission record for that year, not the IMDb record, because the submission title is the authoritative version for eligibility purposes. The database lookup I built for my client took about forty-five minutes total. The manual cross-check against the PDFs for the 2020 through 2024 ceremonies added another two hours because the formatting in those PDFs changes each year. Once the script was stable, adding a new year takes roughly twelve minutes: download the new PDF, run the parser, check for mismatches, and update the SQLite file. If you do not need programmatic access and just want a static table, the Wikipedia page "List of Academy Award winners" is the fastest option. It is well maintained, includes year, category, winner, and year of eligibility, and it updates within forty-eight hours of each ceremony. The downside is that you have to manually verify any figures you publish, since the page is crowd-edited and occasionally lags on correction requests. I caught one error myself where the 2021 Best Picture winner was temporarily mislabeled as The Power of the Dog instead of CODA during a minor vandalism incident, and the page stayed wrong for approximately three hours before an editor reverted it.
What This Data Cannot Tell You
A yearly winner list gives you zero information about margin of victory, runtime, production budget, or box office performance unless you merge in external sources. It also does not capture films that were ineligible because of geographic or language restrictions. Parasite winning Best Picture in 2020 is historically notable precisely because it broke the English-language assumption, but if you only look at the winner list without noting the category split between Best Picture and Best International Feature, you miss that structural context entirely. The dataset also cannot reliably answer questions about streaming-eligible films before the Academy changed its rules. Prior to 2020, a film had to have a theatrical run to qualify. Roger Ebert-era streaming releases do not appear in the historical record for that reason. If your analysis includes the post-2020 period, make sure you adjust for the pandemic-era eligibility extensions, which allowed hybrid and streaming releases to qualify under modified conditions. A dataset that treats 2020 and 2021 winners the same way as 2019 winners will produce skewed conclusions about industry trends. The most practical approach is to treat Oscar Film Winners By Year as a starting point, not a final product. Pull the Kaggle CSV, verify the most recent five years against the official PDFs, merge in your external metadata source if you need budget or box office data, and keep a change log so you can explain discrepancies when someone asks why two sources disagree on a title or a year. That is what actually works in practice.