Wednesday, July 2026

VOL. 19, ISSUE NO. 4 | July 2026

Tech Story

Table 1 from page 74

Abstract

Lightweight multicomponent alloys are of great interest due to the re-markable hardness they exhibit. Optimising the alloying comositions of such alloys to achieve a synergistic combination of weight, strength and ductility remains a challenge. In the present work, a design strategy for lightweight multicomponent alloys is devel-oped to synthesise alloys with im-proved ductility. This novel alloy design strategy combines CALPH-AD-based high-throughput calcula-tions with machine learning (ML) to select suitable compositions. Two new alloy compositions were devel-oped, and their mechanical testing results validated the proposed alloy design procedure.

Introduction

The two most fundamental proper-ties desirable in a structural materi-al are strength and ductility. Additionally, when a component made from such material is intended for non-stationary applications, the weight of the component becomes another critical deciding factor. Therefore, an ideal structural alloy for mobile applications should pos-sess high strength, high ductility, and low density.It is well known that multi-com-ponent alloys, due to the pres-ence of various alloying elements, tend to form multiple phases and solid solutions. Consequently, a multi-component alloy inherent-ly benefits from solid solution and phase-based strengthening mecha-nisms. If such materials could also exhibit higher ductility and have lower density, they would represent a winning combination.For over a century, it has been a con-stant endeavour for physical met-allurgists to develop alloy systems that exhibit all the above properties. This task is non-trivial, as the pos-sible combinations of alloying ele-ments and their resulting properties could be extensive. However, with the advent of artificial intelligence and machine learning, these chal-lenges could be addressed through physics-constrained machine learn-ing approaches.In the current work, an attempt has been made to develop such alloy systems using known and simu-lated physical data within a ma-chine-learning framework.Over the last two decades, alloy discovery activities have been driv-en by computational algorithms. Initially, these activities were pro-pelled by Integrated Computational Materials Engineering (ICME) approaches. In the latter half of the last two decades, they have in-creasingly been driven by machine learning (ML).

Machine Learning (ML) has been extensively applied to the discov-ery of high-entropy alloys (HEAs), a class of materials introduced by Cantor and Yeh in 2004. HEAs typi-cally consist of four or five elements with nearly equal atomic percent-ages, each exceeding five percent. Their high configurational entropy favours the formation of solid-solu-tion phases, predominantly with FCC, BCC, or HCP structures. These alloys are distinguished from con-ventional alloys by their complex compositions, resulting in excep-tional properties such as high me-chanical strength, superior fatigue and wear resistance, excellent mag-netic properties, and exceptional irradiation and corrosion resistance. The inherent complexity and vast performance tuning space of HEAS make traditional experimental tri-al-and-error methods inefficient. and costly. Recently, computer-as-sisted design methods, particu-larly high-throughput calculations (HTC), have significantly advanced HEA research, HTC enables effi-cient, large-scale computations, providing rapid predictions of ma-terial properties and phase stability without relying on empirical param-eters. Despite these advancements, HTC methods require validation through experiments to ensure ac-curate performance predictions, The integration of artificial intelli-gence (Al) and machine learning (ML) further enhances HEA design by capturing complex patterns and predicting properties based oncomprehensive datasets.

The use of ML for alloy discovery has been particularly effective for HEAs due to their unique combi-nation of multiple elements, which prevents reaching a low equilibrium configuration. This results in superi-or chemical, mechanical, and phys-ical properties. ML leverages large datasets to predict phase stability and mechanical properties, moving beyond traditional empirical meth-ods. Various ML algorithms, includ-ing decision trees, neural networks, and support vector machines, mod-el the relationships between alloy composition and properties, with deep learning proving particularly effective.

ML can also address the challeng-es in HEA development by utilising high-throughput techniques (HTCs) and data-driven approaches. HTC and characterisation methods gen-erate extensive datasets, which ML algorithms analyse to uncover in-tricate structure-property relation-ships. This integrated approach, inspired by the Materials Genome Initiative, combines experimental data, theoretical models, and com-putational simulations to predict and optimise HEAs on an unprece dented scale.

ML algorithms enable the rapid identification of optimal alloy.com-positions by modelling the complex interplay between composition and properties. These models screen large compositional spaces and predict desired properties, provid-ing intelligent feedback to guide further experimental efforts. The synergy between high-throughput experiments and ML accelerates the development of advanced HEAS, paving the way for innovative materials with enhanced performance for various applications.

In the present work, the physical properties of interest are ductility and density. While density can be addressed by increasing the num ber of lightweight elements such as aluminium, magnesium, lithium, etc., the studies on ductility have been limited. Ductility, represented by elongation to fracture, has been correlated to Pugh's ratio and Cau-chy pressure. A Pugh's ratio of less than 0.56 or a positive Cauchy pres-sure indicates the element is duc-tile; however, their application to alloys has been limited. Therefore, to achieve high ductility, a com-bination of machine learning and high-throughput CALPHAD is em-ployed, and the results have been experimentally tested.

Design Strategy

The design strategy aims to identify lightweight multicomponent com-positions that can be cast with ac-ceptable ductility. As the number of possible alloy combinations is large and the ability of machine learning models to generalise material be-haviour is limited by available data, the use of CALPHAD in conjunction with ML modelling is proposed to improve the alloy selection process. First, a high-throughput Scheil so-lidification calculation was set up with CALPHAD to explore aluminium-based compositions with favourable phases. The solutions generated were used as input for an ML model to predict ductility (tensileelongation) and corresponding uncertainty. The final alloys selected were synthesised and mechanically tested to validate the design strat with higher reported solubilities in aluminium, namely, zinc, magne-sium, silicon and copper. The prima-ry output from these calculations is FCC volume fractions, however, secondary estimates such as soli-dus temperature, latent heat capac ity, Gibbs free energy and enthalpy are also obtained for each calcula-tion point to examine correlations with ML predictions.

Machine Learning Modelling

The compositions from high-throughput calculations are used in the ML model to predict their ductility. The modelling pro-cess consists of the following steps: data collection, feature extraction, feature screening, model selection, hyperparameter tuning, and ductility predictions.

Chart showing Calculations

Data Collection

Data is essential for training ma-chine learning models, and given PanAl2022 thermodynamic data-base. This solidification model is based on the assumption of cam-plete diffusional mixing in the liq-uid state but no diffusion occurring in the solid state, which is a more accurate representation of solidifi-cation conditions than an equilibri-um solidification model is based on infinite diffusion in both states. The alloying elements chosen are those egy. A flowchart of the proposed approach is given in Fig. 1

Calphad High-throughput Calculations

To identify suitable composition windows with acceptable ductility and low density, a high-through-put Scheil solidification calcula-tion is set up in Pandat using the large enough data, all ML models are reported to perform similarly. In materials science, however, there exists an infinitely large number of alloy combinations with only a lim-ited amount of experimental data available. This limitation in data ne-cessitates the careful selection of ML models.The main target data for our mod-el is ductility, which is represented by elongation exhibited in uniaxial tension/compression testing. We collected experimental data from published literature which includes data for pure alloys and conven-tional alloys as well as HEA alloys. The data labels collected included composition details, primary and secondary processing details, and experimental conditions as well as the target material property. The elements present in the collected data are shown by their frequency of occurrence in Fig. 2

Feature extraction

A large number of features were generated for each composition present in the dataset denoting the physical and chemical attributes of the alloying elements. These fea-tures were created using Python libraries MASTML and Matminer. They include the composition-al mean, max and min of physical properties such as atomic num-ber factor, electrochemical factor, group number factor, cohesion-en-ergy factor, Mendeleev number fac-tor and many more. The total fea-tures generated are 554, and they will be screened to reduce their di-mensionality.

Feature Screening

As the number of features generated in the prior step is very large, a feature screening step is imple-mented to identify the key features relating to material ductility. A two-step feature screening method is adopted. The first step is to remove highly correlated features based on Pear-son's correlation coefficient (r). If two features have a high correlation, then the information they represent is similar, and it would be redundant to include both features. A thresh-old of Ir > 0.7 is selected, resulting in the removal of 433 features with 121 features remaining. For further screening of features, a second step based on Recursive Feature Elimination (RFE) is con-structed. RFE is an iterative meth-od wherein in each iteration an ML model is developed using existing features and the least significant feature is removed. The process continues until a predefined num-ber of features is obtained. In the present study, the RFE algorithm is paired with a Random Forest (RF) model, which provides it with met-rics for ranking feature importance.

Model Selection

As the problem of modelling mate-rial ductility is a regression one with labelled data as an input, a vast va-riety of supervised machine learn-ing algorithms can be used. Four different ML algorithms are used to model ductility, and the model with the highest testing scores is chosen. The models considered are: (a) Linear Regression (LR) (b) Random Forest (RF) (c) Gradient Boosting Regression (XGB) (d) Kernel Ridge Regression (KRR) While a Linear Regression algorithm models the data as having a linear dependence between the indepen-dent and dependent variables, Ran-dom Forests use bagging to build a large collection of de-correlated trees and then average their influ-ence. The random forest predictor can be written mathematically as Eq. 1, where B is a tree generated by bagging and e characterises the bth random forest tree in terms of split variables, cutpoints at each node, and terminal-node values. Gradient Boosting, on the other hand, utilises ensemble methods to build multiple weak models and ag-gregates them to get better perfor-mance as a whole. Kernel Ridge Regression (KRR) combines ridge regression (linear least squares with L2 regularisa-tion) with a kernel trick, which al-lows it to learn non-linear relation-ships between input variables and the target variable. These models are trained using the repeated k-fold cross-validation technique, where the number of re-peats is 3 and the number of folds is 5. In n-fold cross-validation, the to-tal data is randomly divided into n sets, with n-1 sets used for training and the rest used for testing. This testing and training step is repeated n times until all sets have been used for both testing and training. In a repeated n-fold cross-valida-tion technique, the n-fold cross-val-idation procedure is simply repeat-ed m times, with the partitioning of data into the n-folds being different in each of the m repeats due to ran-domisation. The results show that the RF model outperforms all other models. The highest testing R³ score (coefficientof determination) achieved during model selection is 0.602 (averaged across all the folds) with an input of 22 RFE-screened alloy features

Chart showing TECH STORY

To assess the importance of fea-tures, a Pearson's correlation heat-map is generated as shown in Fig. 3a. The highest positive correlation is observed with "Avg deviation in electronegativity", while the high-est negative correlation is with "BCC Fermi difference".

To further improve the RF model, a hyperparameter optimisation loop was set up to optimise the param-eters "n_estimators" and "max_features", followed by a 10-repeat 10-fold cross-validation.

The final training and testing results are shown in Fig. 3(c) & (d). The fi-nal testing R² score of the model is 0.61 averaged across all folds, and the training score is 0.926. The test-ing results indicate that the model. is overpredicting the elongation at lower values and underpredicting it for higher values.

Experimental Procedures

Based on the alloy design strategy, two new lightweight alloys are selected for casting.

The selected alloys are produced by performing induction melting of pure Al, pure Zn and the AlCu50 master alloy, followed by perma-nent mould casting.

The moulds were preheated to 200°C, and the alloys were quenched immediately after casting.

Tensile samples are machined out of the as-cast samples using a wire EDM with a gauge length of 22 mm. Uniaxial tension tests are carried out using an MTS tensile tester of 20 kN capacity at a strain rate of 0.001/s. A laser extensometer was used to measure the strain in the gauge length.

Results and Discussion

The results of the machine learning algorithms to predict the properties of potential alloys are displayed in Fig. 4

The primary data to feed machine learning algorithms for each alloy system was generated using ther-modynamic calculations using CAL-PHAD. Figure 4a depicts the varia-tion and solidity as temperature for the range of algorithms being eval-uated using machine learning and their potential phase content (FCC phase volume fraction) along with predicted elongation and its stan-dard deviation.

As can be seen, there are inherent patterns which relate to phase frac-tion, elongation and standard deviation in elongation.

The set of all variables, both the feature vectors and the output, are plotted as the heat map in Fig. 4b. This shows the relative correlation of each variable.

The set of all variables, both the feature vectors and the output, are plotted as the heat map in Fig. 4b. This shows the relative correlation of each variable.

Therefore, this figure cannot be tak-en as a relation between various variables but should only be taken as a potential set of important features required in the machine learn ing algorithm.To visualise the predicted four-di-mensional space of solidus tem-perature, FCC volume fraction, elongation and elongation standard deviation, it has been projected on a two-dimensional plane containing just FCC volume fraction and elon-gation, as shown in Fig. 4c.

However, a colour encoding based on standard deviation elongation has been used to identify various regions. It is found that despite a lower-di-mensional projection, there is a strong grouping of data based on standard deviation elongation. In general, one would like to have high elongation and low standard devia-tion in elongation, however, as can be seen in the distribution of pre-dicted data, such an instance does not exist.

The choice of these two points was not arbitrary but is based on sound physical metallurgical principles since these two particles had higher FCC volume fractions, their predict ed large elongation has more sound physics-based reasoning apart from the predictions from the machine learning algorithm. These two compositions were: AIS5Cu5Zn40 Al60Cu5Zn35 As described before, these two spe cific alloys were experimentally fab-ricated and tested for their strength and elongation. They both, being high in aluminium content, were in-herently lightweight. The stress-strain curve of the tensile test conducted on these two alloys is provided in Fig. 4(d).

As can be seen, the as-cast alloy shows over 3% strain to failure and a strength of over 400 MPa. This is quite an achievement since the as-cast alloy system exhibits very low strain to failure.The two alloys also show substantial plastic deformation (nonlinear be-haviour beyond elasticity) and have a substantial area under the curve up to failure

A more apple-to-apple comparison would be possible if these alloys were later processed under ther-mo-mechanical treatment and the wrought properties were used as a baseline for comparison. This, of course, is a preliminary study. and as more data, not just from ther-modynamic calculations but also experimental results, become avail-able, machine learning will become more precise and accurate.

It is envisaged that in the future, with advancement in automated experi-mental systems, artificial intelligence algorithms would not only predict possible alloy systems but would also run them on such automatic ex-perimental systems to self-correct their models, leading to a positive feedback system to perform new al-loy discoveries.

References

1. Guo S., Phase selection rules for cast high entropy alloys: An overview. Materials Science and Technology, 31(10), Pages 1223-1230, 2015.

2. Rahul S. Patel, Sharad Bharti-ya, Ravindra D. Gudi, Physics Constrained Learning in Neural Network based Modelling, IF-AC-PapersOnLine, Vol. 55, Issue 7. Pages 79-85, ISSN 2405-8963, 2022.

3. Shunli Zha et al., Machine learn-ing assisted design of high-en-tropy alloys with ultra-high mi-crohardness and unexpected low density, Materials & Design, Vol. 238, 112634, ISSN 0264-1275, 2024.

4. Wang, J., Kwon, H., Kim, H.S. et al, A neural network model for high entropy alloy design. npj Computational Materials, 9, 60, 2023.

5. Liu, S., Yang, C. Machine Learn-ing Design for High-Entropy Alloys: Models and Algorithms. Metals, 14, 235, 2024.

6. Dimiduk, D.M.. Microstruc ture-Property-Design Rela-tionships in the Simulation Era: An Introduction. In: Ghosh, S., Dimiduk, D. (eds.) Computa-tional Methods for Microstruc-ture-Property Relationships. Springer, Boston, MA, 2011.

7. Backman, D.G., Wei, D.Y., Whitis, D.D. et al., ICME at GE: Acceler-ating the insertion of new ma-terials and processes. JOM, 58, 36-41, 2006.

8. M. Hu, Q. Tan, R. Knibbe, M. Xu, B. Jiang, S. Wang, X. Li, M.X. Zhang, Recent applications of machine learning in alloy design: a review. Materials Science and Engineering R: Reports, 155, Ar-ticle 100746, 2023.

9. Cantor B., Chang I.T.H., Knight P., Vincent A.J.B., Microstructur-al development in equiatomic multicomponent alloys. Materi-als Science and Engineering A, 375-377, 213-218, 2004,

10. Feng, R., Zhang, C., Gao, M.C. et al., High-throughput design of high-performance lightweight high-entropy alloys. Nature Communications, 12, 4329, 2021.

11. Li Ruixuan, Xie Lu, Wang William Yi, Liaw Peter K., Zhang Yong, High-Throughput Calculations for High-Entropy Alloys: A Brief Review. Frontiers in Materials, Vol. 7, 2020.

12. Li, S.; Liu, R.; Yan, H.; Li, Z.; Li, Y.; Li, X., Zhang, Y.; Xiong, B. Ма-chine Learning Phase Prediction of Light-Weight High-Entropy Alloys Containing Aluminium, Magnesium, and Lithium. Metals, 14, 400, 2024.

13. Zhu, W., Huo, W., Wang. S. et al., Machine Learning-Based Hard-ness Prediction of High-Entropy Alloys for Laser Additive Manu-facturing. JOM, 75, 5537-5548, 2023.

14. Uttam Bhandari, Md. Rumman Rafi, Congyan Zhang, Shizhong Yang, Yield strength prediction of high-entropy alloys using machine learning, Materials Today Communications, 26, 101871, ISSN 2352-4928, 2021.

15. Soo Young Lee, Seokyeo-ng Byeon, Hyoung Seop Kim, Hyungyu Jin, Seungchul Lee, Deep learning-based phase pre-diction of high-entropy alloys: Optimisation, generation, and explanation, Materials & Design, 197, 109260, ISSN 0264-1275, 2021.

16. Wang, J., Kwon, H., Kim, H.S. et al., A neural network model for high entropy alloy design. npj Computational Materials, 9, 60, 2023.

17. Liu, S., Yang, C. Machine Learn-ing Design for High-Entropy Alloys: Models and Algorithms. Metals, 14(2), 235, 2024.

18. Shomik Verma, Miguel Rivera, David O. Scanlon, Aron Walsh, Machine learned calibrations to high-throughput molecular ex-cited state calculations. Journal of Chemical Physics, 156(13). 2022.

19. Senkov, O.N., Miracle, D.B. Gen-eralisation of intrinsic duc-tile-to-brittle criteria by Pugh and Pettifor for materials with a cubic crystal structure. Scientific Reports, 11, 4531, 2021.

20.S.-L. Chen, S. Daniel, F. Zhang. Y.A. Chang, X.-Y. Yan, F.-Y. Xie, R. Schmid-Fetzer, W.A. Oates, The PANDAT software package and its applications, Calphad, 26(2), Pages 175-188, ISSN 0364-5916. 2002.

21. Michele Banko and Eric Brill, Scaling to Very Very Large Cor-pora for Natural Language Dis-ambiguation. In Proceedings of the 39th Annual Meeting of the Association for Computational Linguistics, Pages 26-33, Tou-louse, France. Association for Computational Linguistics, 2001.

22. Ryan Jacobs et al., The Materials Simulation Toolkit for Machine Learning (MAST-ML): An auto-mated open source toolkit to accelerate data-driven materials research, Computational Mate-rials Science, 176, 109544, ISSN 0927-0256, 2020.

23. Logan Ward et al., Matminer: An open source toolkit for materials data mining, Computational Ma-terials Science, 152, Pages 60-69, ISSN 0927-0256, 2018.

24. Darst, B.F., Malecki, K.C. & Engel-man, C.D. Using recursive feature elimination in random forest to account for correlated variables in high dimensional data. BMC Genetics, 19 (Suppl. 1), 65, 2018.

25.D. Maulud, A.M. Abdulazeez, A Review on Linear Regres-sion Comprehensive in Machine Learning, JASTT, Vol. 1, No. 2, pp. 140-147, 2020.

26. Breiman, L. Random Forests. Machine Learning, 45, 5-32, 200

Table 1 from page 82