So you want to actually use machine learning in geotechnics instead of just putting it on a resume

I've spent the better part of a decade working with soil data, and the first thing you need to understand is that geotechnical datasets are a mess. They are small, they are expensive to collect, and they have more missing values than most people outside this field realize. I once had a project where a borehole log had only 14 samples across 40 meters of stratigraphy, and the client wanted a settlement prediction model. We made it work, but I need to explain how you get there before we talk about tools. This is the application of statistical learning methods to predict soil behavior, classify strata, estimate parameters from indirect measurements, or model foundation performance. The practical implementations you will run into most often are predicting bearing capacity from CPT data, classifying soil types from geophysical logs, estimating shear strength parameters from index properties, and forecasting consolidation settlement. Everything else is mostly research paper noise at this point. The actual workflow starts with understanding what you are trying to predict and whether you even have enough data for it. This is where most people fail. They grab a random forest implementation from GitHub and feed it fifty rows of CPT data without thinking about the fundamental problem. A geotechnical engineer once showed me a model that claimed 97 percent accuracy on predicting liquefaction potential. The dataset had 200 positive cases and 30 negative ones. The model learned to classify everything as liquefiable and achieved 87 percent accuracy. That should tell you something about how machine learning gets used in this field.

Data collection and the problems nobody warns you about

Geotechnical data comes from multiple sources and they do not talk to each other. You have CPTu logs in .csv files with inconsistent column naming, borehole logs in PDFs scanned at low resolution, laboratory test results in spreadsheets that three different engineers formatted differently, and case history records from older projects that exist only as photographs of hand-written field notes. Your first month is almost entirely spent on data reconciliation before any modeling begins. The biggest issue is unit inconsistency. One contractor reports q_c in kPa, another in MPa, and the lab reports Su in both kPa and t/m² depending on which software they used. I had a gradient boosting model degrade from R² of 0.82 to 0.31 because someone had mixed kN/m³ and lb/ft³ in the unit weight column without telling anyone. I wrote a validation script that checks every column against physical bounds for soil mechanics. Any value outside those bounds gets flagged and sent back for manual review. It adds about three hours to a typical preprocessing pipeline but prevents catastrophic errors downstream. Missing data is another reality. Standard imputation methods like mean substitution destroy the covariance structure in your features. If you have a geotechnical dataset where moisture content and plasticity index are correlated, replacing missing values with column means breaks that relationship and biases your predictions. I use iterative imputation with a Gaussian process regressor, which preserves the multivariate structure better than KNN imputation for the small sample sizes you typically encounter in geotechnical work. The tradeoff is that it takes longer to run and needs careful tuning of the iteration count.

Feature selection matters more than algorithm choice

People obsess over whether to use XGBoost, random forest, or neural networks. In my experience, feature engineering and selection account for roughly 70 percent of your model performance. The remaining 30 percent is split between data quality and hyperparameter tuning. Pick the wrong features and even the best algorithm will underperform a simple linear regression with the right inputs. In geotechnical applications, some features are inherently more predictive than others. The normalized cone penetration ratio R_k = (f_c / q_c) * 100 relates friction to total cone resistance and is a much stronger predictor of soil behavior type than either parameter alone. Power law relationships between SPT N-values and CPT tip resistance capture soil stiffness better than raw values. These transformations are not intuitive if you have never worked with this kind of data, and they are worth your time. I encountered a specific edge case that illustrates why feature selection is critical. I was building a model to predict undrained shear strength from CPT data for a soft clay deposit in Bangkok. The initial model using raw q_t, u_1, and I_w features gave disappointing results with cross-validation R² around 0.45. I tried adding polynomial features and interaction terms, which actually made it worse at 0.38. The breakthrough came when I stopped feeding the model raw CPT parameters and instead used stress-normalized indices. The factor of normalized cone penetration N_qt and the dimensionless pore pressure parameter B_q turned out to be the decisive features, pushing cross-validation R² to 0.79. The raw parameters were dominated by overburden stress effects, which masked the soil-specific behavior the model needed to learn.

Get the Full Details

Machine Learning for Data-Centric Geotechnics (Challenges in Geotechnical and Rock Engineering ...
Machine Learning for Data-Centric Geotechnics (Challenges in Geotechnical and Rock Engineering ...

Model selection and the validation approach that actually works

Geotechnical datasets are small, which means standard k-fold cross-validation gives you unreliable performance estimates. With fewer than 200 samples, a single outlier in your training fold can shift your results dramatically. Leave-one-out cross-validation sounds attractive but has high variance. I use repeated stratified k-fold with k=5 repeated ten times, which provides a reasonable balance between bias and variance for typical geotechnical sample sizes. For classification problems like soil behavior type identification, class imbalance is almost always present. Most deposits are predominantly one or two soil types. I handle this with SMOTE-NC, which generates synthetic samples for minority classes while preserving the categorical nature of non-numeric features. Regular SMOTE fails here because interpolating between two soil types might create a physically impossible feature combination. SMOTE-NC respects the feature space structure. Neural networks are not a panacea in geotechnical engineering. They require substantially more data than tree-based methods to generalize, and they tend to overfit quickly on small geotechnical datasets. I have found that gradient boosting machines like XGBoost or LightGBM consistently outperform neural networks for most geotechnical prediction tasks with datasets under 500 samples. The exception is when you have large case history databases with thousands of records, such as settlement predictions from established regional datasets. Even then, tree-based methods usually match or exceed neural network performance with significantly less training time and no need for GPU hardware.

Implementation details and practical workflow

Here is the setup I use most often. Python with scikit-learn for preprocessing and baseline models, XGBoost or LightGBM for the primary classifier or regressor, and SHAP for interpretability. If you are publishing or submitting work to clients, you need to explain your model, and SHAP values give you feature importance rankings that non-technical stakeholders can understand without requiring a statistics degree. The pipeline looks like this: Step one: Load and clean your data. Run validation scripts against physical bounds. Flag anomalies. Document every data source and how you handled missing values. This documentation step is what separates professional work from hobby projects.

Step two: Engineer features. Add stress-normalized indices, derived ratios, and domain-specific transformations. Do not skip this. Raw geotechnical parameters rarely feed well into models without normalization or transformation. Step three: Split your data. Use geographic or stratigraphic splitting when possible rather than random splitting. If your data comes from multiple sites, hold out entire sites for testing. Random splitting inflates performance estimates because samples from the same site share underlying characteristics. Step four: Train and validate. Start with a simple model as a baseline. A logistic regression or linear regression tells you whether your features have predictive power at all. If the baseline performs poorly, no amount of model complexity will fix it.

New paper: Physics-informed machine learning in geotechnical engineering: a direction paper ...
New paper: Physics-informed machine learning in geotechnical engineering: a direction paper ...

Step five: Interpret the results. Check SHAP values, partial dependence plots, and residual distributions. Look for systematic bias in your predictions. If your model consistently underpredicts strength in a certain depth range, that is a signal that you need additional features or that your training data is incomplete in that range.

Where this approach breaks down

Machine learning geotechnical engineering fails when you try to apply a model trained on one soil type or geological setting to a fundamentally different context. I saw a model built on Scandinavian glacial clays applied to Japanese volcanic soils with completely different mineralogy and fabric. The predictions were nonsensical and the engineer who used it nearly specified an unsafe foundation design. Models do not understand physics. They understand patterns in your training data. When those patterns do not exist in your new application domain, the model will confidently produce wrong answers, and you will not know until it is too late. Another limitation is extrapolation. Tree-based models cannot predict values outside the range of their training data. If your CPT data covers q_t values from 2 to 15 MPa and you encounter a layer with q_t of 25 MPa, your model will cap its predictions at the maximum training value. This is particularly dangerous in geotechnical engineering because the most critical designs often involve extreme conditions that may not be represented in your dataset. Always check whether your prediction ranges fall within your training data ranges before using model outputs for design decisions. Uncertainty quantification is also weak in standard implementations. A point prediction of 150 kPa for undrained shear strength does not tell you whether the true value is likely 140 or 200 kPa. For geotechnical design, that difference matters enormously. I use quantile regression forests to generate prediction intervals rather than single point estimates. This adds some computational overhead but provides the uncertainty bounds that engineers actually need for risk assessment.

What I would do differently if starting over

I would invest more time in building a proper feature store and data lineage tracking system from the beginning. The manual reconciliation work I spent weeks on each project could have been reduced to hours if I had standardized data schemas and versioned my datasets properly. I also wish I had learned to use Bayesian optimization for hyperparameter tuning earlier instead of relying on grid search, which wastes computation on regions of the parameter space that are unlikely to improve performance. The most important thing to remember is that machine learning in geotechnical engineering is a tool for augmenting engineering judgment, not replacing it. The models work best when you use them to identify patterns in existing case histories, flag anomalous data points that deserve further investigation, and provide preliminary estimates that you then refine through conventional analysis. They are weakest when you treat them as black boxes that produce design values without questioning whether those values make physical sense. If you want to get started, the Python ecosystem handles this well. The core packages you need are scikit-learn, xgboost, shap, and imbalanced-learn for the resampling techniques. I typically start from a template repository that already has the data validation pipeline and cross-validation setup configured, which saves me several days of boilerplate work on each new project.

PolyU project promotes multidimensional machine learning in geotechnical engineering supported ...
PolyU project promotes multidimensional machine learning in geotechnical engineering supported ...