Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Insights into Cation Ordering of Double Perovskite Oxides from Machine Learning and Causal Relations

This work investigates origins of cation ordering in double perovskites using first-principles theory computations combined with machine learning (ML) and causal relations. We have considered various oxidation states of A, A', B, and B' from the family of transition metal ions to construct a diverse compositional space. A conventional framework employing traditional ML classification algorithms such as Random Forest (RF) coupled with appropriate features including geometry-driven and key structural modes leads to accurate prediction (~98%) of A-site cation ordering. We have evaluated the accuracy of ML models by employing analyses of decision paths, assignments of probabilistic confidence bound, and finally a direct non-Gaussian acyclic structural equation model to investigate causality. Our study suggests that structural modes are crucial for classifying layered, columnar, and rock-salt ordering. The charge difference between A and A' is the most important feature for predicting clear layered ordering, which in turn depends on the B and B' charge separation. We have also designed mathematical relationships with these features to derive energy differences to form clear layered ordering. Here, the trilinear coupling between tilt, in-phase rotation, and A-site antiferroelectric displacement in the Landau free-energy expansion becomes the necessary condition behind formation of A-site cation ordering.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An Observational Comparison of Level of Neutral Buoyancy and Level of Maximum Detrainment in Tropical Deep Convective Clouds

Tropical deep convective clouds are important drivers of large-scale atmospheric circulation representing the main vertical transport pathway through the depth of the troposphere for heat, momentum, water, and chemical species. The strength and depth of this transport are impacted by the convective updraft size and intensity that are driven by buoyancy, dynamical forcing, and mixing of environmental air, i.e., entrainment. In this paper, we identify tropical deep convective systems with well-defined forward anvils using Atmospheric Radiation Measurement (ARM) ground-based profiling radars, at three ARM fixed-sites in the Tropical Western Pacific (TWP; i.e., Manus, Nauru, Darwin) and three ARM Mobile Facility deployments in Niamey, Niger; Gan Island, Maldives; and Manacapuru, Brazil. We use the difference between the level of neutral buoyancy (LNB) and the level of maximum detrainment (LMD) as a proxy for the effective bulk convective entrainment (εproxy). The LNB, the theoretical height that a parcel raised above the level of free convection would reach with no mixing, is calculated based on pre-convection radiosonde measurements using parcel theory. The LMD is the height of the maximum reflectivity observed in forward anvil clouds by profiling radars. Deep convective systems over the TWP show higher LNBs that extend to 16.3 km on average and larger εproxy (median LNB minus LMD up to 6.5 km) compared to their continental counterparts in the Amazon and West Africa. Oceanic conditions show larger convective available potential energy (CAPE) coupled with higher moisture at low levels which favors larger εproxy. In contrast, continental cases initiate and develop under high convective inhibition, steeper environmental lapse rate, and high wind shear conditions, which show smaller offset between LNB and LMD. Deep convective cases that promote significant cold pools at the surface experience less εproxy. Using a Random Forest regression algorithm, CAPE is associated with the highest feature importance score for predicting convective εproxy, followed by low-level relative humidity. For continental cases, the low-level wind shear also indicates higher importance.

54 ENVIRONMENTAL SCIENCES↗

A comparison of machine learning methods to classify radioactive elements using prompt-gamma-ray neutron activation data

The detection of illicit radiological materials is critical to establishing a robust second line of defence in nuclear security. Neutron-capture prompt-gamma activation analysis (PGAA) can be used to detect multiple radioactive materials across the entire Periodic Table. However, long detection times and a high rate of false positives pose a significant hindrance in the deployment of PGAA-based systems to identify the presence of illicit substances in nuclear forensics. In the present work, six different machine-learning algorithms were developed to classify radioactive elements based on the PGAA energy spectra. The model performance was evaluated using standard classification metrics and trend curves with an emphasis on comparing the effectiveness of algorithms that are best suited for classifying imbalanced datasets. We analyse the classification performance based on Precision, Recall, F1-score, Specificity, Confusion matrix, ROC-AUC curves, and Geometric Mean Score (GMS) measures. The tree-based algorithms (Decision Trees, Random Forest and AdaBoost) have consistently outperformed Support Vector Machine and K-Nearest Neighbours. Based on the results presented, AdaBoost is the preferred classifier to analyse data containing PGAA spectral information due to the high recall and minimal false negatives reported in the minority class.

97 MATHEMATICS AND COMPUTING↗

Identification of novel organic polar materials: A machine learning study with importance sampling

Recent advances in the synthesis of polar molecular materials have produced practical alternatives to ferroelectric ceramics, opening up exciting new avenues for their incorporation into modern electronic devices. However, in order to realize the full potential of polar polymer and molecular crystals for modern technological applications, it is paramount to assemble and evaluate all the available data for such compounds, identifying descriptors that could be associated with an emergence of ferroelectricity. In this paper, we utilized data-driven approaches to judiciously shortlist candidate materials from a wide chemical space that could possess ferroelectric functionalities. A machine learning study with importance sampling was employed to address the challenge of having a limited amount of available data on already-known organic ferroelectrics. Sets of molecular- and crystal-level descriptors were combined with a Random Forest Regression algorithm in order to predict the spontaneous polarization of the shortlisted compounds. First-principles simulations were performed to further validate the predictions obtained from the machine learning model.

36 MATERIALS SCIENCE↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

Improved sensitivity of the DRIFT-IId directional dark matter experiment using machine learning

We demonstrate a new type of analysis for the DRIFT-IId directional dark matter detector using a machine learning algorithm called a Random Forest Classifier. The analysis labels events as signal or background based on a series of selection parameters, rather than solely applying hard cuts. The analysis efficiency is shown to be comparable to our previous result at high energy but with increased efficiency at lower energies. This leads to a projected sensitivity enhancement of one order of magnitude below a WIMP mass of 15 GeV c -2 and a projected sensitivity limit that reaches down to a WIMP mass of 9 GeV c -2 , which is a first for a directionally sensitive dark matter detector.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine Learning in Infectious Disease for Risk Factor Identification and Hypothesis Generation: Proof of Concept Using Invasive Candidiasis

Machine learning (ML) models can handle large data sets without assuming underlying relationships and can be useful for evaluating disease characteristics, yet they are more commonly used for predicting individual disease risk than for identifying factors at the population level. We offer a proof of concept applying random forest (RF) algorithms to Candida-positive hospital encounters in an electronic health record database of patients in the United States. Candida-positive encounters were extracted from the Cerner HealthFacts database; invasive infections were laboratory-positive sterile site Candida infections. Features included demographics, admission source, care setting, physician specialty, diagnostic and procedure codes, and medications received before the first positive Candida culture. We used RF to assess risk factors for 3 outcomes: any invasive candidiasis (IC) vs non-IC, within-species IC vs non-IC (eg, invasive C. glabrata vs noninvasive C. glabrata), and between-species IC (eg, invasive C. glabrata vs all other IC). Fourteen of 169 (8%) variables were consistently identified as important features in the ML models. When evaluating within-species IC, for example, invasive C. glabrata vs non-invasive C. glabrata, we identified known features like central venous catheters, intensive care unit stay, and gastrointestinal operations. In contrast, important variables for invasive C. glabrata vs all other IC included renal disease and medications like diabetes therapeutics, cholesterol medications, and antiarrhythmics. Known and novel risk factors for IC were identified using ML, demonstrating the hypothesis-generating utility of this approach for infectious disease conditions about which less is known, specifically at the species level or for rarer diseases.

60 APPLIED LIFE SCIENCES↗

Efficient Estimation of Pauli Observables by Derandomization

In this work, we consider the problem of jointly estimating expectation values of many Pauli observables, a crucial subroutine in variational quantum algorithms. Starting with randomized measurements, we propose an efficient derandomization procedure that iteratively replaces random single-qubit measurements by fixed Pauli measurements; the resulting deterministic measurement procedure is guaranteed to perform at least as well as the randomized one. In particular, for estimating any L low-weight Pauli observables, a deterministic measurement on only of order log(L) copies of a quantum state suffices. In some cases, for example, when some of the Pauli observables have high weight, the derandomized procedure is substantially better than the randomized one. Specifically, numerical experiments highlight the advantages of our derandomized protocol over various previous methods for estimating the ground-state energies of small molecules.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Distributionally Safe Path Planning: Wasserstein Safe RRT

In this paper, we propose a Wasserstein metric-based random path planning algorithm. Wasserstein Safe RRT (W-Safe RRT) provides finite-sample probabilistic guarantees on the safety of a returned path in an uncertain obstacle environment. Vehicle and obstacle states are modeled as distributions based upon state and model observations. Additionally, we define limits on distributional sampling error so the Wasserstein distance between a vehicle state distribution and obstacle distributions can be bounded. This enables the algorithm to return safe paths with a confidence bound through combining finite sampling error bounds with calculations of the Wasserstein distance between discrete distributions. W-Safe RRT is compared against a baseline minimum encompassing ball algorithm, which ensures balls that minimally encompass discrete state and obstacle distributions do not overlap. The improved performance is verified in a 3D environment using single, multi, and rotating non-convex obstacle cases, with and without forced obstacle error in adversarial directions, showing that W-Safe RRT can handle poorly modeled complex environments.

42 ENGINEERING↗

Accelerating large scale de novo metagenome assembly using GPUs

Metagenomic workflows involve studying uncultured microorganisms directly from the environment. These environmental samples when processed by modern sequencing machines yield large and complex datasets that exceed the capabilities of metagenomic software. The increasing sizes and complexities of datasets make a strong case for exascale-capable metagenome assemblers. However, the underlying algorithmic motifs are not well suited for GPUs. This poses a challenge since the majority of next-generation supercomputers will rely primarily on GPUs for computation. In this paper we present the first of its kind GPU-Accelerated implementation of the local assembly approach that is an integral part of a widely used large-scale metagenome assembler, MetaHipMer. Local assembly uses algorithms that induce random memory accesses and non-deterministic workloads, which make GPU offloading a challenging task. Our GPU implementation outperforms the CPU version by about 7x and boosts the performance of MetaHipMer by 42% when running on 64 Summit nodes.

Awan, Muaaz Gul↗

ARENA: Adversary-Resistant Evolving Neural Architectures

Neural networks are becoming the cornerstone for national security prediction tasks. However, designing them requires significant research and trial/error, as they have many hyperparameters, including their computation graph (“architecture”). Neural architecture search (NAS) employs secondary optimizers to search for architectures maximizing objectives like accuracy. Evolutionary algorithms (EAs) are the most used class of optimizer for NAS. However, existing Python libraries for writing EAs limit the complexity of experiments a user can design. In this project, we built ARENA, a Python framework that encodes complex, hyper-realistic EAs. ARENA collects detailed information as it runs and is flexible enough to encode non-EA search algorithms. We tested ARENA on 4 toy optimization problems by encoding 3 search algorithms for each—random search, an EA, and simulated annealing. We also designed an EA that performs NAS on the MNIST dataset. Our experiments suggest the potential for immediate mission impact through solving lab-wide optimization problems.

97 MATHEMATICS AND COMPUTING↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

A computational pipeline to generate a synthetic dataset of metal ion sorption to oxides for AI/ML exploration

The charged mineral/electrolyte interfaces are ubiquitous in the surface and subsurface–including the surroundings of the geological disposal sites for radioactive waste. Therefore, understanding how ions interact with charged surfaces is critically important for predicting radionuclide mobility in the case of waste leakage. At present, the Surface Complexation Models (SCMs) are the most successful thermodynamic frameworks to describe ion retention by mineral surfaces. SCMs are interfacial speciation models that account for the effect of the electric field generated by charged surfaces on sorption equilibria. These models have been successfully used to analyze and interpret a broad range of experimental observations including potentiometric and electrokinetic titrations or spectroscopy. Unfortunately, many of the current procedures to solve and fit SCM to experimental data are not optimal, which leads to a non-transferable or non-unique description of interfacial electrostatics and consequently of the strength and extent of ion retention by mineral surfaces. Recent developments in Artificial Intelligence (AI) offer a new avenue to replace SCM solvers and fitting algorithms with trained AI surrogates. Unfortunately, there is a lack of a standardized dataset covering a wide range of SCM parameter values available for AI exploration and training–a gap filled by this study. Here, we described the computational pipeline to generate synthetic SCM data and discussed approaches to transform this dataset into AI-learnable input. First, we used this pipeline to generate a synthetic dataset of electrostatic properties for a broad range of the prototypical oxide/electrolyte interfaces. The next step is to extend this dataset to include complex radionuclide sorption and complexation, and finally, to provide trained AI architectures able to infer SCMs parameter values rapidly from experimental data. Here, we illustrated the AI-surrogate development using the ensemble learning algorithms, such as Random Forest and Gradient Boosting. These surrogate models allow a rapid prediction of the SCM model parameters, do not rely on an initial guess, and guarantee convergence in all cases.

Li, Chunhui↗

Predictive Modeling of NOx Emissions from Lean Direct Injection of Hydrogen and Hydrogen/Natural Gas Blends Using Flame Imaging and Machine Learning

This research paper explores the use of machine learning to relate images of flame structure and luminosity to measured NOx emissions. Images of reactions produced by 16 aero-engine derived injectors for a ground-based turbine operated on a range of fuel compositions, air pressure drops, preheat temperatures and adiabatic flame temperatures were captured and postprocessed. The experimental investigations were conducted under atmospheric conditions, capturing CO, NO and NOx emissions data and OH* chemiluminescence images from 27 test conditions. The injector geometry and test conditions were based on a statistically designed test plan. These results were first analyzed using the traditional analysis approach of analysis of variance (ANOVA). The statistically based test plan yielded 432 data points, leading to a correlation for NOx emissions as a function of injector geometry, test conditions and imaging responses, with 70.2% accuracy. As an alternative approach to predicting emissions using imaging diagnostics as well as injector geometry and test conditions, a random forest machine learning algorithm was also applied to the data and was able to achieve an accuracy of 82.6%. This study offers insights into the factors influencing emissions in ground-based turbines while emphasizing the potential of machine learning algorithms in constructing predictive models for complex systems.

08 HYDROGEN↗

Bursty channel errors and the Viterbi decoder

Recent applications have developed for spread spectrum communications, hardware data transfer, high rate digital systems, etc. that use channels for which errors tend to occur in short bursts in addition to those at random, i.e., compound channels. Viterbi decoding algorithms are generally very good for random error channels but are not as efficient for burst errors or for compound channels. This paper presents the results of a computer simulation study of the performance of various Viterbi decoders when receiving data corrupted with burst and random errors on the same channel. Simulations were performed using hard-decision CPSK.

Ingels, F.↗

Cometary atmospheres: Modeling the spatial distribution of observed neutral radicals

An algorithm for the random walk problem of multiple elastic collisions between newly formed non-thermal neutral cometary radicals and the outflowing cometary molecules was incorporated into the Monte Carlo particle-trajectory model. Preliminary model analysis has shown that the effects of collision on the observed spatial distribution of cometary radicals becomes important for the larger bright comets, especially at smaller values of the helicocentric distance. The model and early results are discussed herein.

Combi, M. R.↗

Application of genetic algorithms to tuning fuzzy control systems

Real number genetic algorithms (GA) were applied for tuning fuzzy membership functions of three controller applications. The first application is our 'Fuzzy Pong' demonstration, a controller that controls a very responsive system. The performance of the automatically tuned membership functions exceeded that of manually tuned membership functions both when the algorithm started with randomly generated functions and with the best manually-tuned functions. The second GA tunes input membership functions to achieve a specified control surface. The third application is a practical one, a motor controller for a printed circuit manufacturing system. The GA alters the positions and overlaps of the membership functions to accomplish the tuning. The applications, the real number GA approach, the fitness function and population parameters, and the performance improvements achieved are discussed. Directions for further research in tuning input and output membership functions and in tuning fuzzy rules are described.

Espy, Todd↗

Local design optimization for composite transport fuselage crown panels

Composite transport fuselage crown panel design and manufacturing plans were optimized to have projected cost and weight savings of 18 and 45 percent, respectively. These savings are close to those quoted as overall NASA Advanced Composite Technology (ACT) program goals. Three local optimization tasks were found to influence the cost and weight of fuselage crown panels. The effects are summarized of each task and the task associated with a design cost model is described in detail. Studies were performed to evaluate the relationship between manufacturing cost and design details. A design tool was developed to aid in these studies. The development of the design tool included combining cost and performance constraints with a random search optimization algorithm. The resulting software was used in a series of optimization studies that evaluated the sensitivity of design variables, guidelines, criteria, and material selection on cost. The effect of blending adjacent design points in a full scale panel subjected to changing load distributions and local variations was shown to be important. Technical issues and directions for future work were identified.

Swanson, G. D.↗