Understanding the Connolly and Begg Database Textbook
Most people pick up Database System By Thomas Connolly because their professor told them to, or because they are studying for a certification and it keeps coming up in reference lists. It is a heavy textbook. The latest editions run past 800 pages and they cover everything from the relational model through to NoSQL hybrids and cloud-based deployment models. That breadth is both its strength and its weakness. It is not a quick reference manual. If you want something you can open at your desk and flip through while troubleshooting a production issue, this is not it. This book is designed for a semester-long course or for self-study over several months. The theory sections are rigorous. The normalization chapters alone could take a week to fully work through if you are actually doing the exercises rather than skimming them. I spent about three weeks working through the normalization material in the earlier editions when I first encountered it back in 2008. I was building internal tools for a logistics company and kept running into update anomalies in what I thought were properly designed tables. The book walks through 1NF through BCNF with enough worked examples that the concepts actually stick. I remember specifically struggling with the difference between 3NF and BCNF on chapter exercises involving multi-valued dependencies. The explanation in the book is clearer than most online resources I found at the time.
What the book covers in practice
The core of the text is the relational model, ER diagramming, and SQL. That covers roughly the first half. Then it moves into advanced design topics like functional dependencies, normalization theory, and query optimization. The second half gets into database administration, distributed databases, and newer topics like data warehousing and NoSQL systems in the later editions. One thing beginners miss about this book is that the SQL chapters assume you already understand the theoretical foundation. If you jump straight into the SQL examples without reading the relational model chapters first, you will find yourself memorizing syntax without understanding why certain constraints exist. I saw this repeatedly with junior developers I mentored early in my career. They could write a join query but had no idea why their database was producing duplicate rows or why their insert statements were failing on referential integrity violations.
Practical problems I ran into using this as a guide
When I worked through the case studies in the book, I tried applying the normalization techniques to a real inventory tracking system I was building. The problem was that the book's examples use clean, textbook data. Real business data is messy. I had a supplier table where the same supplier appeared under slightly different naming conventions because data had been entered by different people over two years. The normalization theory in Connolly and Begg does not address data quality issues like that directly. My workaround was to run a fuzzy matching script using Levenshtein distance to identify near-duplicate supplier records before attempting any normalization. I wrote a quick Python script that flagged records with a similarity score above 0.85, then I manually reviewed the flagged pairs. Only after cleaning the data did the normalization process produce sensible results. The book assumes clean input. In production, that assumption rarely holds.
Get the Full Details
Common pitfalls when studying from this text
The biggest issue I see is that people treat the ER modeling chapters as purely academic. The book teaches you to draw entity-relationship diagrams with perfect cardinality notation, but real systems have requirements that change constantly. I worked on a project where the ER diagram designed following the book's methodology had to be completely revised after three months because a new regulatory requirement forced a structural change that the original design could not accommodate without significant denormalization. The methodology is sound for stable requirements. It is less useful when requirements shift mid-project. Another pitfall is the SQL examples. They use standard SQL syntax that maps closely to Oracle and PostgreSQL, but if you are working in MySQL or SQL Server, some of the advanced features like recursive CTEs or window functions may behave differently or not be available depending on your version. The book notes this in passing but does not dedicate sections to vendor-specific differences.
Counter-intuitive things the book gets right
One insight that surprised me on first reading is the treatment of denormalization. The book does not present denormalization as inherently bad, which many introductory courses imply. It explains that controlled denormalization is sometimes the correct engineering decision when read performance matters more than write performance, and when the data access patterns are predictable and stable. I applied this principle when optimizing a reporting database that was taking too long to generate daily summaries. A small amount of deliberate redundancy in the summary tables cut query times from forty seconds to under two seconds. The book also handles the topic of database triggers in a nuanced way. Rather than simply warning against them, it explains specific scenarios where triggers are appropriate, such as maintaining audit logs or enforcing complex business rules that cannot be expressed through foreign key constraints alone. I have seen triggers misused so badly that they created cascading update loops taking down entire databases. The book teaches you to identify those risk scenarios before writing trigger code.
Limitations of this book
It is expensive. New editions run around sixty to eighty dollars depending on where you buy it. The PDF versions exist in various forms online but I am not going to link to any of them. There are older editions available secondhand for much less, and the core relational database material has not changed significantly between editions. The later editions add more content on NoSQL and cloud databases, so if you only care about traditional relational systems, an older edition is perfectly adequate and costs a fraction of the price. The book also has limited coverage of modern ORM frameworks and application-level database interaction patterns. If you are a developer who mostly works with Entity Framework, Django ORM, or similar tools, you will find the raw SQL chapters useful but the book does not bridge the gap between raw SQL and ORM-generated queries. That is a gap you need to fill from other sources. For a more practical companion, I would recommend pairing it with something like SQL Server or PostgreSQL documentation depending on your platform, along with a hands-on project that forces you to apply the normalization concepts to actual messy data. Reading the book alone will not make you proficient at database design. You need to build something, break it, and fix it using the principles the text describes.
Where to find the book
You can order it from major retailers like Amazon, or check your university bookstore if you are a student. Cengage, the publisher, sometimes offers digital rental options that are cheaper than buying the full eBook. The ISBN varies by edition, so make sure you get the latest one your course or project requires. The 14th edition is the most recent as of my current knowledge. The book is worth the time investment if you are serious about database design. It will not make you a database administrator overnight, and it will not replace hands-on experience with actual database engines. But the theoretical foundation it provides is solid, and it remains one of the more respected textbooks in the field for that reason.