What You Actually Need in Your Python Stack
I keep seeing people ask for a complete Libraries List For Data Science and then try to install everything at once. That is a fast way to create environment conflicts that will waste your weekend. Here is how I actually structure mine and which packages earn their place on disk. At the bottom of everything sits NumPy. It handles multi-dimensional arrays and basic linear algebra. If you are doing anything numerical, you are going to need this whether you admit it or not. It is the foundation that most other packages depend on, so it goes in first. Then there is pandas. This is where you load CSVs, clean messy data, merge tables, and do the boring work that takes up 70 percent of a data science project. I used to skip the documentation and just figure it out by trial and error. That changed after I spent three hours debugging a merge that failed because two columns had different index types. pandas handles index alignment silently, which is convenient until it silently does the wrong thing. Always check your dtypes after a merge. I run df.dtypes.head(20) as a habit now.
scikit-learn is your workhorse for machine learning. Classification, regression, clustering, preprocessing, model selection. It does not do deep learning, and that is intentional. The API is consistent across all estimators, which means once you learn fit and predict, you know the interface for everything in the library. I have written custom transformers that plug right into pipelines without breaking the standard methods. That consistency is worth more than any single feature.
Data Visualization That Does Not Waste Time
Matplotlib is the default and it is everywhere. You will see it in tutorials, papers, and stack overflow answers. It is also verbose and ugly by default if you do not configure it. I set up a rcParams block in my startup file that changes the figure size, font, and grid settings once and forget about it. The alternative is typing the same twenty lines of formatting code into every notebook. Seaborn sits on top of matplotlib and makes statistical plots reasonable without a lot of effort. Heatmaps, pair plots, distribution charts. It defaults to looking acceptable, which saves you from spending time on aesthetics when you could be doing something else. The tradeoff is that you lose some control over individual elements compared to raw matplotlib. Plotly is worth adding if you need interactive visualizations or dashboards. The export to static images works fine, but the real value shows up when you are presenting to stakeholders who want to zoom and hover. I used it for a quarterly report and cut the back-and-forth about chart details by half because people could explore the data themselves instead of asking me to regenerate plots.
When You Need More Power
XGBoost and LightGBM are the two gradient boosting frameworks that matter for tabular data. They handle missing values differently, which matters more than people admit. XGBoost treats missing as a separate category by default, while LightGBM sends missing values down a specific leaf during training. I ran into this when switching between them on a Kaggle competition. My validation score dropped because the missing data pattern in my test set was being handled inconsistently. The fix was setting handle_missing explicitly in both and verifying the behavior on a small subset first. TensorFlow and PyTorch are the deep learning options. PyTorch is easier to debug because the computation graph is dynamic. TensorFlow has better production deployment tools through SavedModel and TensorFlow Serving. I chose PyTorch for research and prototyping because I can print intermediate tensors without wrapping everything in a tf.function. For deployment, I package the model and use TorchServe or export to ONNX if the pipeline requires it. SciPy handles specialized statistics and scientific computing. Sparse matrices, integration, optimization. Most people never touch it directly because pandas and scikit-learn abstract the common cases, but when you need a statistical test that is not in sklearn or a sparse matrix operation that is not in scipy.sparse, you go here. I needed a custom kernel for Gaussian process regression once and ended up writing it against scipy.linalg functions. It took longer than expected because the documentation assumes you already know linear algebra internals.
Tooling That Makes Everything Less Painful
Jupyter Lab or VS Code with Jupyter extensions. The notebook environment is still the most common workspace for exploratory analysis, even though I have moved most production code into scripts. Notebooks are useful for iteration and bad at version control. I keep notebooks for exploration and move anything that needs to run repeatedly into Python modules. DBT if you are working with SQL-heavy workflows. It has a steep learning curve but pays off quickly if your data lives in a warehouse and you need reliable transformations. I replaced a shell script that ran hourly queries with a dbt project in about two days. The model testing and documentation features caught three bugs that the old script would have silently produced. Airflow or Prefect for orchestration. Airflow is heavier and more established. Prefect is simpler and easier to debug locally. I ran an Airflow DAG once where a task failed silently because the XCom size limit was exceeded. The error message pointed at something completely unrelated. Prefect would have shown me the actual problem in the UI. I switched after that.
What I Leave Out and Why
There are a hundred other packages people add to their Lists. LangChain for LLM workflows, spaCy for NLP, CV2 for computer vision. These are not wrong additions, but they belong in separate environments. Mixing them all into one base installation creates dependency hell faster than almost anything else. I keep a base environment for the core packages and create isolated ones for specialized work. Statsmodels is another one that sits between pandas and scikit-learn. It gives you proper statistical inference with p-values and confidence intervals. Scikit-learn intentionally does not provide these. If your work requires hypothesis testing or you need to report statistical significance, statsmodels is the right tool. The API is less uniform than sklearn, which is the cost of being more statistically rigorous. Dask for out-of-core computation when data exceeds memory. It parallelizes pandas and NumPy operations with minimal code changes. The catch is that it adds overhead and does not help with everything. Single-threaded operations on small datasets run slower with Dask than with vanilla pandas. I only reach for it when I hit memory limits or need to parallelize across many files.
Setting up a reliable environment takes about fifteen minutes the first time if you use conda or a virtual environment manager properly. After that, you are spending most of your time installing niche packages for specific projects and dealing with incompatibilities between them. The Libraries List For Data Science is not a fixed thing. It grows slowly and sheds packages as your needs change. The ones I listed above are the ones I have not removed in three years.