So You Want to Bake Your Models
The whole "baking" pipeline starts with a raw trained model, usually something like a .pt or .pb file, and the goal is to produce something that runs fast on your target hardware without losing meaningful accuracy. It sounds simple, but the intermediate steps are where most people waste a day or two before realizing they set something wrong in the middle. Here's the actual order I follow, not the theoretical one. First, pick your target runtime. I can't stress this enough because people routinely bake for CPU, then realize their deployment target is an NVIDIA Jetson Orin, and now they need to rework everything. Pick the runtime before you do any graph surgery. The main options are TensorRT for NVIDIA GPU, Core ML for Apple silicon, TFLite for mobile CPU/NPU, and ONNX as a middle ground that sometimes works and sometimes doesn't depending on which operators your model uses. Once the target is locked, export the model to an intermediate format if your framework doesn't natively support the runtime. PyTorch models become ONNX with torch.onnx.export. TensorFlow 2.x models can export directly to SavedModel or Keras HDF5. This step alone takes about 30 seconds to 5 minutes depending on model size, and the export log will tell you if any operators failed to convert. I once had a custom LayerNorm implementation silently fall back to a Python op during ONNX export, which produced a valid .onnx file that worked on CPU and completely crashed on GPU. The workaround was enabling the onnxruntime extended diagnostics flag and comparing the operator table before and after export against the original TensorFlow graph.
After a clean export, run a benchmark on the unoptimized model first. Get baseline latency and memory usage. Then apply optimization passes. Quantization is the big one — INT8 reduces model size by roughly four times compared to FP32 and typically cuts inference latency by 30-60 percent on supported hardware. The catch is that pure post-training quantization loses accuracy on sensitive models. The fix is quantization-aware training or using TF Lite's full integer quantization with a calibration dataset. I ran into this exact problem with a YOLOv8 variant where PTQ dropped mAP from 47.2 to 41.8, but calibration with 200 representative images brought it back to 46.1 with no retraining needed. Pruning comes next if you have compute budget to spare during the build process. Channel pruning removes entire convolutional filters based on L1 norm thresholds, which usually yields another 20-40 percent speedup on top of quantization for standard architectures. Structured pruning requires your target runtime to support the sparsity format natively, which most don't well except for NVIDIA's CSR format in TensorRT. You'll also want to fuse convolution-BN pairs if they haven't been fused already. Every unfused Conv+BN pair adds an extra kernel launch and redundant memory ops that add up fast on edge devices. The actual baking step — whether that's trtexec for TensorRT, coremltools compile for Core ML, or tflite_convert for TFLite — produces the final artifact. Verify it by running the same validation dataset you used for baseline measurement. If accuracy drops more than 1-2 percent from the original, something went wrong in the optimization chain and you need to isolate which pass caused it by rolling back one step at a time.
What Nobody Tells You About the Process
Graph optimization is usually the highest-leverage step and the one people skip because the tools hide it. A lot of operators that look expensive in isolation — elementwise additions, scalar broadcasts, reshape operations — get completely eliminated or fused during the graph optimization phase. Running your model through Netron after export helps you see what actually survived the fuser pass versus what was theoretically in your original architecture. Memory profiling matters more than latency numbers early on. I've seen people chase 2ms latency improvements only to discover the baked model OOMs on the target device because the intermediate activation tensors weren't being reused properly. TensorRT's memory planning can be tuned with --memPoolSize, and setting the workspace size too aggressively can cause the builder to fall back to slower algorithms or fail entirely. A safe starting point is 512MB for small models and 2-4GB for vision transformers, but you should measure actual peak memory during a real inference run before committing to a workspace size. There's also a real limitation to keep in mind: baking doesn't fix a bad model. If your original architecture is inefficient or over-parameterized, baking just makes the inefficiency run faster. For production deployments, I'd recommend doing a lightweight architecture review before investing time in the baking pipeline. Sometimes a smaller model with proper baking outperforms a larger unbaked one on the same hardware.
Get the Full Details

The whole pipeline from raw model to deployed artifact typically takes between 2 and 6 hours for a first pass, depending on how many optimization iterations you need. The biggest time sink is always the accuracy-debug cycle after quantization or pruning, not the baking tooling itself. Budget accordingly.