About this project

# Physics-informed ML for Hall–Petch strengthening in FCC MPEAs This repository serves as the companion material for the *Acta Materialia* manuscript titled "Revisiting Hall–Petch strengthening in FCC multi-principal element alloys: what grain size, composition, and processing can and cannot explain" by M. Mulukutla, S. P. Padhy, et al. It contains the complete dataset, the full analysis pipeline, every result table and figure presented in the paper, LaTeX sources, a comprehensive report regenerated from live results, and regression tests that lock the manuscript's canonical values. ## The Question Machine-learning property models can mislead alloy design when interpolation is mistaken for transfer, or when a physically motivated descriptor is read as a mechanism without showing that it adds information beyond the measured inputs. This repository works that problem end-to-end on the yield strength (YS) and Vickers hardness (HV) of 94 FCC alloy conditions in the Al–Co–Cr–Cu–Fe–Mn–Ni–V system. ## Five Model Families Analyses are organised by the five families of the manuscript. A higher family number means more flexibility, not more physical fidelity. - **Family 1**: Classical Hall–Petch (classical and alternative grain-size laws) - **Family 2**: Physics descriptors (VLC, Labusch, Toda-Caraballo; Wen; PCA-OLS) - **Family 3**: Composition / processing (the M-model hierarchy, incl. M15) - **Family 4**: Non-linear ML (matched linear and non-linear estimators) - **Family 5**: Symbolic regression (PySR, SISSO, fixed closed forms) Validation protocol includes 5-fold, LOO, LOBO, literature, and singularity checks. ## Headline Results All values are pooled out-of-fold Q², recomputed at test time from `data/derived/data_with_vlc.csv`. | Target | Model | 5-fold | LOO | LOBO | |---|---|---|---|---| | YS | Family 1 classical Hall–Petch | 0.405 | 0.406 | 0.373 | | YS | Family 2 PCA-OLS (fold-contained) | 0.462 | 0.480 | 0.362 | | YS | Family 3 M3 (composition-dependent σ₀) | 0.666 | 0.652 | 0.625 | | YS | **Family 3 M15 (+ SD_grain interaction)** | **0.731** | **0.694** | **0.694** | | YS | Family 4 Lasso S2 (fixed settings) | 0.683 | 0.674 | 0.621 | | YS | Family 4 LightGBM S2 (fixed settings) | 0.632 | 0.574 | 0.615 | | YS | Family 4 Bayesian ridge S2 (nested ARMOTE-CV) | 0.685 | 0.666 | 0.613 | | HV | Family 1 classical Hall–Petch | 0.086 | 0.136 | −0.077 | | HV | Family 4 LightGBM S1 | 0.195 | 0.323 | 0.112 | | HV | Family 5 fixed form (post-selected) | 0.725 | 0.727 | 0.636 | Three findings do most of the work: - **Grain-distribution width helps only in a specific form.** Adding SD_grain to M3 as an independent additive term raises LOO to 0.668 but lowers LOBO to 0.595. Letting it modify the Hall–Petch term instead (M15) gives 0.694 at both, with ΔBIC = −20.1 against M3. The additive control is what makes that distinction visible. - **Non-linearity buys nothing at matched inputs.** Lasso and LightGBM reach 0.621 and 0.615 under batch-held-out validation with overlapping bootstrap intervals. - **Nothing transfers to the literature.** Every evaluated closed form has negative R² on the direct-measurement literature records, so that set is reported as a stress test rather than an external benchmark. ## Layout ``` data/raw/ Grain_Size_Summary_v3.xlsx — the only true input data/derived/ computed descriptors, regenerated by `make data` scripts/ _config.py shared paths; makes every family folder importable _figstyle.py one palette and one type scale for every figure 00_data_preparation/ descriptors, VLC quantities, feature ladder 01_family1_grain_size/ scaling laws, within-replicate slopes 02_family2_physics_descriptors/ SSS benchmark, redundancy audit, PCA-OLS 03_family3_composition_processing/ M-model hierarchy, SD_grain models 04_family4_nonlinear_ml/ tuned panel, matched-input comparison 05_family5_symbolic_regression/ PySR grid, SISSO variants 06_validation/ grouped CV, literature test, singularity audit 07_hardness_tabor/ Tabor ratio and HV–YS rank analysis figures/ every publication figure results/ CSV outputs; each is produced by a named script analysis_plots/ exploratory and diagnostic figures paper/ main.tex, supplementary.tex, references.bib, figures/ report/ generate_report.py + Comprehensive_Analysis_Report.docx notebook/ generator + generated .ipynb docs/ reproducing.md, validation_protocol.md tests/ regression tests locking the canonical values ``` ## Quick start ```bash pip install -r requirements.txt make test # ~1 s — recomputes and checks every headline value make figures # regenerate all publication figures from cached results make report # rebuild the Word report from live results make paper # build main.pdf and supplementary.pdf make help # every target ``` Full stage order and timings: `docs/reproducing.md`. ## Reproducibility `make test` recomputes the M-model hierarchy, the Family 1 baselines, the Tabor ratio and the dataset audit **from the raw derived data** and asserts them against the values printed in the manuscript. It also checks that every figure and table is cross-referenced in the text and that every included figure file exists. If an analysis changes, the tests fail before the manuscript can drift. The **nested ARMOTE-CV panel** is verified rather than re-run. Its generator stores one Optuna study, one selected parameter set, and one pair of feature and target scalers per fold. Across the six LOBO folds the five tuned Bayesian-ridge hyperparameters take six distinct values, and each scaler is fitted on its training split alone. Reloading those objects and predicting each held-out batch reproduces all six fold scores exactly and gives a pooled LOBO Q² of 0.613 (RMSE 50.5 MPa). The nesting is therefore demonstrated, not asserted. Those artifacts ship with the repository, under `scripts/04_family4_nonlinear_ml/armote_cv/` — see the README there for what was kept and what was left upstream. Two analyses are **archival** and are not regenerated here, in the manuscript, or by `make`: | Archival result | Why | |---|---| | Bayesian PSIS-LOO stacking weights | PyMC posterior draws were not retained | | Outer-loop PySR performance | equations were selected on complete-data fronts | The manuscript reports these as archival, and the corresponding frequentist quantities that *can* be recomputed (M15's OLS coefficients, information criteria and grouped-CV scores) are regenerated by `make sdgrain`. ## Figures Every figure imports `scripts/_figstyle.py`, so the whole document uses one Okabe–Ito palette and one type scale. Colour is never the only channel: batches carry distinct marker shapes, model families are separated by position, and signed quantities print their value. Figures are drawn at their final printed width so a point size in the script is the same point size on the page. The palette was checked under simulated deuteranopia, protanopia and tritanopia. ## Citation See `CITATION.cff`. Archival metadata for Zenodo is in `.zenodo.json`. ## Licence MIT — see `LICENSE`. The raw experimental dataset and the literature compilation are included; cited-literature PDFs are not redistributed. ## Acknowledgment Sponsored by the Army Research Laboratory under Cooperative Agreement W911NF-22-2-0106.