Getting Started with Candy Box Project Math
Candy Box Project Math is a lightweight mathematical modeling framework used for inventory forecasting, demand clustering, and batch optimization in small-scale production environments. It was built to handle situations where you have limited SKUs but complex seasonal patterns and variable lead times. Most people who come across it are either running a small warehouse operation or trying to model something similar to one. At its foundation, the system treats your product set as a distributed probability problem rather than a deterministic one. You input your order history, seasonality flags, and supplier lead time windows, and it generates a distribution map for each SKU. From there, you can calculate reorder points, safety stock buffers, and batch sizes without spinning up a full ERP system. The math itself relies on a modified Poisson-Gamma conjugate model with a seasonal adjustment layer. That sounds more complicated than it is in practice. You load your data, pick a confidence level (0.90 to 0.98 is standard), and the framework spits out recommended order quantities. I spent about three hours debugging my first implementation because I kept feeding it daily data when the system was designed for weekly aggregation. That matters. The model smooths over short-term noise, and daily snapshots create artificial spikes that throw off the whole distribution.
How It Actually Works in Practice
Here's the straightforward version of getting started. You download the repository from the project's main GitHub page, clone it into your working directory, and run the dependency install script. It requires Python 3.9 or later, numpy, and a few other standard packages. The setup takes roughly ten minutes on a normal machine. Once installed, you structure your input CSV with at least five columns: SKU identifier, date, units sold, unit cost, and lead time in days. A sixth column for promotional flags is optional but useful. The validation script will catch most structural issues before you even run the core model. I'd recommend running validation first every single time. I learned that after wasting an afternoon on a silent failure caused by a date format mismatch that the downstream process swallowed without error.
Candy Box Project Math Configuration Options
The config file lives in the root directory and uses YAML format. There are about a dozen parameters you can tweak, but only four actually matter for most use cases: the confidence_interval, the seasonal_period, the min_history_months, and the buffer_mode. Set confidence_interval to 0.95 unless you have a specific reason not to. Seasonal_period should match your business cycle in weeks. Most operations I've seen sitting at either 4, 12, or 52 depending on whether they're looking at monthly, quarterly, or yearly patterns. Min_history_months defaults to six but you can go lower if your data is sparse. Buffer_mode controls how aggressively the model adds safety stock. Conservative mode adds a larger margin around the predicted range. Aggressive mode trims it closer to the lower bound. Pick conservative if your lead times are unpredictable. Aggressive works fine if you have reliable suppliers and fast shipping. One thing beginners miss is the outlier handling parameter. By default the model clips anything beyond three standard deviations from the mean. That's reasonable for most cases, but if you're dealing with products that genuinely have lumpy demand patterns, that clip will understate your risk. I ran into this with a hardware store client who stocked seasonal items like patio heaters and space heaters. Their demand wasn't noisy, it was naturally bimodal. The default clipping masked half of their actual peak periods and they ended up understocking by roughly forty percent during the relevant seasons. Setting outlier_clipping to false and switching to a wider confidence interval fixed the problem.
Get the Full Details

Running the Model and Interpreting Results
After your data is validated and your config is set, you run the main execution script. It processes each SKU independently and outputs a JSON report along with a CSV summary. The JSON contains the full probability distributions for each item. The CSV is the shorter version with just the recommended order quantities, reorder points, and suggested buffer levels. Most people stop at the CSV, which is usually enough for day-to-day operations. The output includes a forecast horizon measured in weeks. By default it covers the next twelve weeks. You can extend this to twenty-four weeks if you need longer visibility, but accuracy degrades noticeably past week sixteen. That's a limitation of the framework itself, not a bug in your setup. The underlying model loses resolution over longer windows because it has to extrapolate seasonal patterns further out. If you need longer-range visibility, you're better off running the model in rolling monthly batches instead.
Common Pitfalls to Avoid
First, don't mix product categories in the same dataset without flagging them separately. The model treats every row as independent, but if your SKUs span wildly different price points and demand velocities, the aggregated statistics can be misleading. Keep electronics, clothing, and consumables in separate datasets or add a category flag and run per-category configs. Second, watch your lead time column. If you enter average lead times instead of distribution data, the model will underestimate variance. The ideal approach is to input lead time as a list of historical values per SKU, not a single number. The framework supports that format. A lot of users skip it because it takes more effort to gather, but the results are meaningfully worse without it. I once modeled for a client who used a flat seven-day lead time across all suppliers. When one supplier suddenly pushed out to fourteen days during a supply disruption, the model had no mechanism to flag the increased risk because it never learned the variance in the first place. Third, the framework doesn't handle stockouts in the historical input gracefully. If a SKU ran out of inventory for multiple periods in your data, the sales records will show zero demand during those weeks, and the model will treat that as legitimate zero demand rather than a constrained signal. This is a real issue and not a minor quirk. You need to mark stockout periods explicitly in a separate column and include them in the input. The validation script will warn you about zero-demand gaps longer than two weeks, but it won't auto-correct them.
Where the Framework Falls Short
Candy Box Project Math isn't designed for high-volume real-time operations. It's a batch-processing tool. You run it on a schedule, usually weekly or biweekly, and apply the recommendations manually or push them into your purchasing workflow. It won't integrate natively with most modern ERP systems out of the box. You'll need to write a small connector or use the CSV export for manual imports. The framework also assumes stationary demand patterns with seasonal overlays. If your product category is shifting rapidly due to market trends, new competition, or changing consumer behavior, the model will lag behind reality. It can't adapt quickly to structural breaks. I'd recommend supplementing it with a lightweight trend adjustment layer if your category experiences more than fifteen percent year-over-year change in demand volume. For teams that need multi-echelon inventory optimization, cross-warehouse allocation, or dynamic pricing integration, this framework will hit its ceiling pretty fast. Those problems require different tools. Candy Box Project Math is best suited for single-warehouse operations with under two hundred active SKUs and a planning horizon of one to six months.

Where to Get It
The project is hosted on GitHub under the name candy-box-project-math. The README contains the full installation guide and example datasets. There's also a small community Discord server for troubleshooting. Most of the regular contributors hang out there on weekdays. Issues tend to get answered within a day or two unless they involve edge cases that require source-level debugging, which can take longer.