Understanding the Sutro Size Guide for Practical Deployment
The Sutro Size Guide exists to help you pick the right model variant from the Sutro lineup, and honestly, most people skip reading it and then wonder why their latency numbers are terrible or their bill is absurd. The guide breaks down the available sizes by parameter count, context length, inference speed, and recommended workload. That sounds straightforward until you actually try to match those specs to your use case. I spent a couple weeks last year trying to migrate a customer support routing pipeline from a smaller Sutro variant to the largest one available, mainly because I assumed more parameters meant better comprehension. It did not work the way I expected. The larger model handled nuance fine, but our p99 latency jumped from roughly 800 milliseconds to nearly four seconds, and we had to provision twice as many GPU instances to maintain the same throughput. The Size Guide mentions compute requirements, but it does not emphasize how much of a scheduling nightmare that becomes when you are working with shared inference clusters.
Sutro Size Guide breakdown
The Size Guide organizes models into at least three tiers: a small/efficient variant optimized for high-throughput tasks like classification or short-form generation, a medium variant that balances latency and reasoning capacity, and a larger variant aimed at complex reasoning, long-context comprehension, and multi-step tasks. Each tier has a listed parameter range, maximum context window, and typical token-per-second throughput under standard batching. The guide also flags which hardware each size targets, usually split between CPU-favorable and GPU-required deployments. The part everyone glosses over is the context-length tiers within each size. A medium model might offer 8K, 32K, and 128K context options, and switching between them does not just change memory usage linearly. Attention computation scales differently depending on the kernel and whether you are using flash attention or a less optimized fallback. I learned that the hard way when a client enabled 128K context on a medium Sutro instance expecting similar throughput to their 8K setup. Their inference cost tripled and throughput dropped by about sixty percent. The workaround was dropping back to 32K for most requests and only using 128K on a small subset of inputs that actually required it, which brought costs back to acceptable levels while preserving the edge-case coverage.
How to read the guide without making costly mistakes
Start with your latency budget, not your accuracy expectations. That is backwards from what most engineers do, but it matters more than you think. A model that answers correctly but takes five seconds to respond will get disabled in production faster than a slightly less accurate one that responds in under a second. Check your maximum acceptable response time first, then filter the Size Guide to variants that can meet it at your target batch size. Next, look at your input profile. If most of your requests are under 500 tokens, running a large context model is almost always a waste. I have seen teams do this repeatedly. The small variant on Sutro handles short inputs efficiently, and the Size Guide lists throughput numbers that make this obvious if you actually read them. The medium variant is the sweet spot for inputs between one and eight thousand tokens. The large variant starts making sense when you are feeding it five-plus thousand token prompts with multi-step reasoning requirements, or when you need the full context window for document analysis. Cost modeling is the step most people skip until after deployment. The Size Guide provides per-token pricing estimates, but those are baseline figures. Real cost depends on your batching strategy, your idle instance management, and whether you are using spot or reserved capacity. I once calculated that a small Sutro instance running with aggressive batching and early-exit optimizations was half the price of a medium instance running naively, even though the medium instance was handling a wider range of tasks. The lesson was that the Size Guide gives you the map, but you still have to drive carefully.
Get the Full Details

When the guide is wrong for your situation
There are edge cases where following the Size Guide directly leads to poor outcomes. One common scenario involves multilingual workloads. Certain Sutro variants perform significantly better on non-English languages at smaller sizes because the training mix was heavier on those languages, while the largest variants lean toward English-centric reasoning benchmarks. If your pipeline processes mostly Spanish or Japanese inputs, the medium or even small variant may outperform the largest one on your actual metrics. The guide does not always surface this distinction prominently. Another scenario is fine-tuning. If you plan to adapt a Sutro model to your domain with supervised fine-tuning or reinforcement learning, starting with a smaller variant is often more practical. Fine-tuning larger models requires substantially more GPU memory, longer training runs, and more careful learning-rate scheduling. I fine-tuned a medium Sutro model on a proprietary legal-document dataset and found that the small variant, after fine-tuning, matched the unfine-tuned medium model on our evaluation set while running at twice the speed. The Size Guide would have pointed me toward the medium or large variant based on raw capability numbers, but the fine-tuned small variant was the actual better choice.
A practical workflow for picking the right size
Run a benchmark before committing. Take your actual production data, split it into representative samples, and test at least two adjacent tiers from the Size Guide. Measure latency, throughput, token cost, and task accuracy. Do this on the same hardware you plan to deploy on. Cloud environments can behave very differently from on-prem setups, and even different cloud regions can vary in available instance types. I typically run three to five representative requests through each candidate model, then scale up to a hundred for throughput measurements. This usually takes about twenty minutes for a small dataset and gives me enough signal to make a decision. I never rely solely on the Size Guide numbers or vendor benchmarks, because those are measured under ideal conditions that rarely match real traffic patterns. After deployment, monitor your actual performance against the guide's estimates for at least a week. I have caught cases where the production environment had memory pressure from co-located services, which degraded inference performance below what the guide promised. Adjusting the instance type or reducing concurrent connections resolved it without needing to upgrade the model size.
What the Sutro Size Guide does not tell you
It does not cover dynamic routing, which is a technique where you send simple requests to a smaller model and only escalate complex ones to a larger model. This can cut costs by thirty to fifty percent in mixed-workload environments, and it requires logic outside the Size Guide's scope. If your traffic has a wide variance in request complexity, building a routing layer is worth the engineering effort. It also does not address model versioning differences within the same size tier. Sutro releases updates periodically, and a newer medium variant may outperform an older large variant on certain tasks due to training data improvements or architecture changes. Always check the release notes alongside the Size Guide rather than assuming size is the only relevant variable. The Sutro Size Guide is a solid starting point, but it is not a replacement for testing your own workload. The numbers in the guide are accurate under standard conditions, and the tier descriptions are well-organized, but the gap between those numbers and production reality is where most projects go off track. Pay attention to your latency constraints, your actual input distributions, and your cost ceiling, and you will land on the right size without unnecessary over-provisioning or regrettable downgrades.