How to Actually Prepare for Deep Learning Technical Interviews
I've sat on both sides of these interviews for years. Most candidates come in knowing the surface-level definitions but fall apart the moment you push past textbook answers. This guide covers what actually comes up and how to prepare without wasting time on stuff that won't help. Let me start with a scenario that keeps coming back. A few years ago I interviewed someone who could explain transformer attention from first principles, derive the scaled dot-product formula, and draw the full encoder-decoder architecture. Then I asked: "Your model's validation loss starts plateauing but the training loss is still dropping. Walk me through how you'd debug this." They froze. Completely. Could recite dropout and L2 regularization but couldn't talk through an actual debugging strategy. That's the gap most candidates have. Theory is memorized. Practice is empty. Here's what I actually ask, and what a solid answer looks like.
Backpropagation questions are always on the list. Not "explain backprop" — that's too basic. More like: "Walk me through the gradient flow for a residual block and explain what changes when you add batch normalization before and after the activation." The difference between a mediocre and a strong answer here is understanding that skip connections don't just shortcut gradients — they change the effective depth of the network, which is why very deep networks with residuals train reliably while plain deep networks vanish. If someone says "batch norm stabilizes training" and stops there, that's not enough. They should mention the running mean/variance computation during inference, the epsilon term, and why you don't backprop through the normalization statistics the same way you backprop through the affine parameters. Optimization questions tend to separate people who've actually trained models from people who've read papers. I'll ask about Adam versus SGD with momentum. The common wrong answer is that Adam is always better. It's not. Adam converges faster in early training but frequently generalizes worse on vision tasks. The reason involves the per-parameter adaptive learning rates accumulating history in a way that can lock parameters into suboptimal regions. SGD with momentum, despite looking crude, often finds sharper minima that generalize better because the uniform learning rate prevents that premature lock-in. I look for candidates who can discuss the empirical finding, the theoretical explanation, and why AdamW fixed part of the issue by decoupling weight decay from the adaptive learning rate. Here's another one I use regularly: "Explain how you'd implement mixed precision training and what specific numerical issues you'd need to watch for." The surface answer is fp16 saves memory and speeds up compute. The real answer involves gradient scaling to prevent underflow, the master weights technique in optimizers, and why you still want fp32 for certain operations like reduction kernels. I once had a candidate who implemented AMP correctly but didn't understand why loss scaling needed to be dynamic rather than static. When I pressed them on it, they couldn't explain the overflow detection logic. That's the level of detail that matters.
Distributed training comes up more than you'd think, even for non-ML-platform roles. Data parallelism versus model parallelism. Gradient all-reduce. Pipeline parallelism. ZeRO stages. A candidate should understand why communication becomes the bottleneck before compute does in most distributed setups, and how techniques like gradient compression and overlapping communication with computation address that. I once worked with someone who tried to implement pipeline parallelism without accounting for bubble time and wondered why their scaling was terrible. The fix was inserting lookahead scheduling, but they didn't know the term for it. Knowing the vocabulary helps, but understanding the constraint that creates the problem is what separates people who've struggled with this from people who've read about it. For architecture questions, I avoid the generic "compare CNN and Transformer" prompt. Instead I ask something like: "You're building a model to classify satellite imagery at multiple resolutions. Would you use a pure CNN, a pure Transformer, or a hybrid? What are the actual tradeoffs and when would each fail?" A good answer discusses inductive bias — CNNs encode translation equivariance which satellite images benefit from at multiple scales — versus the data efficiency advantage of Transformers when you have enough samples. The failure mode for CNNs here is handling long-range dependencies between distant regions of the image. The failure mode for Transformers is quadratic attention cost across high-resolution patches, which is why Swin and similar architectures exist. Regularization and generalization questions are where people tend to bluff. I'll ask about label smoothing, stochastic depth, mixup, and cutmix. The expected answer goes beyond naming techniques. Label smoothing prevents overconfident predictions which helps calibration and generalization. Stochastic depth drops entire layers during training as a form of ensemble regularization. Mixup and cutmix create smoother decision boundaries by interpolating between examples. But the nuanced part is knowing when each breaks. Mixup can hurt when classes are already well-separated and the interpolation creates ambiguous examples that don't reflect the true data distribution. Cutmix on small objects can crop out the signal entirely. I've seen production models degrade because someone applied mixup to a dataset with many rare classes without adjusting the interpolation coefficient.
Get the Full Details

Here's a practical tip that most candidates miss. When you're asked to design a model from scratch in an interview, start by stating your constraints. "What's the input size? What's the latency budget? How much training data do we have? What's the class distribution like?" The best designs I've seen came from people who asked clarifying questions first. A candidate who immediately starts drawing layers without understanding the problem space is signaling that they've only ever followed tutorials. I once rejected someone who designed a 50-layer ResNet for a problem that needed a 3-layer MLP because they were so focused on showing off architecture knowledge that they ignored the actual requirements. Simple models with the right inductive bias beat complex models every time when the data is limited. One more topic that separates experienced practitioners: evaluation and debugging. I'll throw a curveball like "your model achieves 94% accuracy on training and 62% on validation. What do you check first?" The instinctive answer is overfitting, but that's too shallow. You need to check for data leakage first — is the validation set actually held out? Then check class imbalance. Then check whether the validation distribution matches the training distribution. I've seen production systems fail because the training data came from one season and the deployment data came from another. The model learned seasonal patterns, not the underlying signal. Data augmentation can't fix distribution shift. You need domain adaptation techniques or retraining with representative data. Understanding the failure modes of your tools matters more than knowing how to use them. PyTorch's autograd is convenient until you hit a memory wall. TensorFlow's eager execution is great until you need graph-level optimizations. The frameworks are tools, not solutions. I recommend practicing explanations out loud to someone who knows nothing about deep learning. If you can't explain attention mechanisms to a smart non-technical person, you don't understand them well enough. The Feynman technique isn't just a study hack — it's the single best preparation method for these interviews because it exposes gaps in your understanding that rereading papers never will.
Preparation timeline: two weeks is realistic if you're already familiar with the material. Four weeks if you're starting from scratch. Focus on understanding over memorization. Interviewers can tell when someone has rehearsed an answer versus when someone actually understands the concept. The former falls apart under follow-up questions. The latter adapts.