Why I Keep Coming Back to This Book
I've taught enough data science courses to know that most textbooks either read like a dictionary or skip from math straight to production code without explaining the bridge. John Kelleher's textbook avoids both traps by actually showing you what the math looks like when it's doing something, not just when it's sitting on a whiteboard. The book is freely available under a Creative Commons license, and if you've ever been frustrated by other intro texts, you should probably look at it first before buying anything else. The full title is Data Science, and it's structured around three parts: thinking about data, representing and transforming it, and applying algorithms. That third part is where most learners either connect the dots or drop out, depending on how much hand-holding they got earlier. Kelleher covers Python implementations for most topics, which matters because a lot of free resources give you pseudocode or R and leave you figuring out the translation yourself. Here's what the practical setup looks like. You download the free PDF or read it online. The companion Python code lives on his website. I found the repository layout straightforward but not perfectly organized — the example scripts are functional, though. When I was using this alongside a graduate class back in 2018, the chapter on decision trees had a small bug in the entropy calculation in one of the sample scripts. It wasn't a conceptual error in the book, just a typo in the code. I worked around it by writing a quick helper function that computed entropy from scratch instead of trusting the provided version. If you're running through the examples yourself and your output doesn't match the text, check for that kind of thing first before assuming the theory is wrong.
What beginners consistently miss about this book is that Kelleher deliberately includes chapters on data handling and visualization before he gets to any machine learning. That pacing isn't filler. Most people who try to jump straight into model building come back later and realize their feature engineering was garbage because they never learned how to systematically explore a dataset. The book forces that discipline on you. It's annoying in the moment if you're impatient, but it prevents entire categories of failures down the line. Another thing that's not obvious from reading the table of contents: the treatment of unsupervised learning is surprisingly practical. Most introductory books treat clustering as an afterthought. Kelleher actually walks through real cluster validation techniques, which is something I wish more people understood before they started feeding unlabeled data into k-means and calling it analysis. I've seen projects derail completely because someone didn't know how to check whether their clusters were meaningful. The book's section on this is probably the most underappreciated part. There are real limitations. The Python examples use a version of scikit-learn that's a few years old now. Some API calls will throw warnings in current releases. It's not a dealbreaker — the conceptual mapping still holds — but you might spend time updating function signatures rather than learning new material. The deep learning coverage is also light. If you want neural networks, you'll need another resource. The book introduces the basics, but it's not trying to be comprehensive on that front.
The book is available free at datasciencebook.ie. No registration wall, no hidden paywall, no "sign up for our newsletter to get the link." Just download it. If you're approaching data science from a non-technical background, the early chapters on thinking with data are worth reading even if you plan to skip the code later. If you already know some programming, the progression from data wrangling through to evaluation metrics gives you a framework most tutorials don't provide because they're solving specific problems instead of teaching how to think about them. I've recommended this resource to people with zero coding experience and people with five years in the field who just needed a solid refresher. It works for both, though you'll skim differently. The depth is consistent throughout, which is unusual for a book at this scope. Most authors either over-explain the easy parts and rush the hard ones, or vice versa. Kelleher keeps the pacing reasonable across the board, and that's not something to take for granted.
Get the Full Details
