What You Need to Know Before Digging Into PGA Tournament Archives
The PGA has a long, complicated record, and most of what you'll find online is either a Wikipedia summary or a sponsored article from a betting site. If you actually want to dig into match results, season statistics, and historical bracket data, you need to know where the real records live and how to read them. The official PGA website maintains a database, but it's not exactly intuitive. I spent about three weeks last fall building a scraper that pulled shot-by-shot data from 1998 to 2023 for a personal project, and even with that kind of access, some of the gaps are maddening. The PGA of America runs the main championships, while the PGA Tour operates separately and handles the weekly tour events. That separation matters because they don't always share data cleanly. The official PGA of America site covers the Ryder Cup, the PGA Championship, and the Palmer Cup. The PGA Tour site covers the FedEx Cup standings, tournament brackets, and player stats. When you're looking at pre-1990 results, both sites get spotty, and the third-party statistical databases like GolfStat and StatGolf end up being your best bet. I ran into a real issue last year trying to verify a claim about Jack Nicklaus's major championship round splits. The PGA of America site listed his scores by round, but the FedExCup.com archive had slightly different figures for two tournaments in the 1970s. The discrepancy came down to a different scoring convention for match play versus stroke play being merged into the same statistical table. I cross-referenced the original newspaper archives from the Times-Picayune and the LA Times, and the newspaper numbers matched the PGA of America version. The lesson here is that you should never trust a single source for pre-1990 data. Verify against at least two independent records before using anything in a publication or analysis.
How to Pull Historical Tournament Data
If you're building a dataset, start with the PGA Tour's public JSON endpoints. They expose historical leaderboards, round-by-round splits, and money list data. The API isn't documented publicly, but the network tab in any browser dev tools will show you the calls. A typical request looks like pulling a tournament page and extracting the leaderboard JSON, which comes back with player IDs, round scores, totals, and positions. It takes about twenty minutes to set up a proper scraper once you figure out the endpoint structure. For the older stuff, before the 1990s, you're mostly working with scanned PDFs and JPEGs of paper scorecards. The PGA of America has digitized some of their archives, but the quality varies wildly. Some tournaments are clear scans, others are nearly illegible. I ended up spending about forty hours manually transcribing 1974 PGA Championship data because the existing digital versions had inconsistent formatting. If you're doing this kind of work, budget two to three weeks for a single tournament's worth of pre-1980 records. There are third-party packages that handle the scraping for you. On GitHub, there's a project called pgatour-data that wraps the unofficial API and gives you a pandas DataFrame with season data. It's not officially maintained anymore, but the core functions still work for anything through 2023. The trade-off is that you won't get the very latest rounds if the official site changes its structure, which it does roughly every two years without warning.
Common Pitfalls That Waste People's Time
The biggest problem people run into is conflating the PGA of America with the PGA Tour. They are separate organizations with separate histories. When you see a result labeled "PGA Championship," that's run by the PGA of America. When you see a regular PGA Tour event, that's a different body. Mixing them up ruins any dataset you're building. I've seen this mistake in at least half the articles I've read about golf statistics. Another issue is score verification. Some historical tournaments have disputed scores because of weather delays, course changes, or scorecard errors that were corrected after the fact. The 1969 Open Championship at Turnberry is a famous example where several players' final rounds were adjusted after play concluded. If you're compiling historical accuracy, note any known corrections and cite the source of the correction rather than presenting the adjusted numbers as the original recorded scores. Data completeness is also a real problem. Round-by-round scoring for PGA Tour events didn't become consistently tracked until the mid-1980s. Before that, most records only show final tournament totals and money earned. If you need granular performance data for earlier eras, you'll be working with far less information, and the gaps are permanent. There's no workaround for that. You just have to acknowledge what isn't available and design your analysis around the data you actually have.
Get the Full Details

The official PGA history pages are reasonably thorough for post-1980 events. For anything before that, you're better off checking the USGA archives for major championships and the individual tournament archives where they exist. Some clubs like Oakmont and Merion maintain their own detailed records going back to the early 1900s, and those are often more complete than the national organization's pages. If you just need quick reference data without building your own system, the PGA Tour app and website have searchable tournament results going back several decades. It's not as clean as a spreadsheet, but it's faster than hunting down individual scorecards for one-off lookups.