Setting Up a Real Working Environment for Data Science in VS

Most people jump into this with defaults and wonder why everything breaks. I spent about three weeks last year debugging a notebook that refused to load pandas past version 2.0 because I hadn't properly separated my base Python from my project environments. Here's how to actually do it without losing your mind.

Getting Started with Vs Data Science

Start with VS Code, not Jupyter alone. The standalone Jupyter interface is fine for quick exploration but it falls apart the moment you need version control, multi-file projects, or CI/CD pipelines. Install the Python extension by Microsoft and the Jupyter extension. That's it for the base. Don't install everything they recommend in the marketplace sidebar unless you know what each one does. Create a virtual environment immediately. Do not use conda as your primary setup if you can avoid it — pip and venv or, better yet, uv are faster and less prone to dependency conflicts. I run mine like this:

uv venv .venv --python 3.11 source .venv/bin/activate uv pip install jupyter pandas numpy scikit-learn matplotlib seaborn

Using uv instead of pip cuts environment creation time from about 45 seconds down to roughly 8 seconds on a typical machine. The difference is noticeable when you're spinning up environments for five or six different projects in a week.

Project Structure That Actually Works

A proper layout matters more than people admit. I used to dump everything into notebooks until a colleague pointed out that my model training code was tangled with my data cleaning, visualization, and export logic in a single 400-line cell block. It was unmaintainable. Now I structure things like this: - src/models/ for model definitions - src/data/ for loading and preprocessing functions - notebooks/ for exploratory work only - configs/ for hyperparameters and experiment metadata - data/raw and data/processed kept separate with a .gitignore that excludes both The key insight most beginners miss is that notebooks should never be your final deliverable. They're for exploration. Any code you want to reproduce needs to live in .py files. I wrote a small adapter that lets me import from src/ directly into notebooks, and it saves me from copy-pasting cells between files when I'm iterating.

The Jupyter Kernel Problem

Here's something the documentation glosses over. When you open a notebook, VS Code doesn't always connect to the kernel you think it's using. I had a project where my imports were silently falling back to the global Python installation instead of the virtual environment I'd activated. The symptom was subtle — scipy worked but scikit-learn had a completely different version than what pip showed. You catch this by running `python -c "import sys; print(sys.executable)"` in your first notebook cell. If that path doesn't include your virtual environment directory, you've got a mismatch. Select your kernel explicitly through the bottom-right corner of VS Code or the command palette. Make sure it points to the Python executable inside your .venv.

Practical Debugging: A Specific Case

Last month I hit a problem where GPU memory allocations were failing during a training run, but only in VS Code's integrated Jupyter environment. The same script ran fine from a terminal. The issue turned out to be that VS Code's Python extension was pre-loading certain CUDA-related packages into the kernel process before my notebook even executed, and this consumed about 800MB of GPU memory before training started. I confirmed it by checking nvidia-smi right after the kernel connected but before running any cells. The workaround was straightforward — I added a notebook setting in the workspace config:

"jupyter.notebookFileRoot": "${workspaceFolder}", "jupyter.pythonCompletionPreferredInterpreter": "${workspaceFolder}/.venv/bin/python", "python.analysis.extraPaths": ["${workspaceFolder}/src"]

And I created a custom launch configuration that disables the auto-import behavior. Training runs dropped from crashing at epoch two to completing in about nine minutes on an A10G.

What This Approach Doesn't Handle Well

This setup isn't for everyone. If you're doing collaborative teaching or working with stakeholders who need to run notebooks without touching the terminal, VS Code adds friction. JupyterLab in a browser is more accessible for that crowd. Also, if your project relies heavily on R or Julia, the experience in VS Code is good but not as polished as the native IDEs for those languages. For pure Python data science workflows, though, this is about as efficient as it gets.

Essential Extensions Worth Keeping

I recommend sticking to a minimal set. The ones I use regularly: Python, Jupyter, Pylance, Flake8 or Ruff for linting, and remote-ssh if you're connecting to a GPU server. Everything else is noise. I've seen people install a dozen extensions and then complain about slow startup times. VS Code handles about 5 to 6 well-chosen extensions fine. Past that, you're paying a performance penalty for features you rarely touch. For the data-specific workflow, Ruff has replaced both Flake8 and autopep8 for me. It formats and lints in under a second on files that used to take five seconds each.