Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Denitrification and the challenge of scaling microsite knowledge to the globe

Here, our knowledge of microbial processes—who is responsible for what, the rates at which they occur, and the substrates consumed and products produced—is imperfect for many if not most taxa, but even less is known about how microsite processes scale to the ecosystem and thence the globe. In both natural and managed environments, scaling links fundamental knowledge to application and also allows for global assessments of the importance of microbial processes. But rarely is scaling straightforward: More often than not, process rates in situ are distributed in a highly skewed fashion, under the influence of multiple interacting controls, and thus often difficult to sample, quantify, and predict. To date, quantitative models of many important processes fail to capture daily, seasonal, and annual fluxes with the precision needed to effect meaningful management outcomes. Nitrogen cycle processes are a case in point, and denitrification is a prime example. Statistical models based on machine learning can improve predictability and identify the best environmental predictors but are—by themselves—insufficient for revealing process-level knowledge gaps or predicting outcomes under novel environmental conditions. Hybrid models that incorporate well-calibrated process models as predictors for machine learning algorithms can provide both improved understanding and more reliable forecasts under environmental conditions not yet experienced. Incorporating trait-based models into such efforts promises to improve predictions and understanding still further, but much more development is needed.

59 BASIC BIOLOGICAL SCIENCES↗

Quantifying leaf symptoms of sorghum charcoal rot in images of field‐grown plants using deep neural networks

Abstract Charcoal rot of sorghum (CRS) is a significant disease affecting sorghum crops, with limited genetic resistance available. The causative agent, Macrophomina phaseolina (Tassi) Goid, is a highly destructive fungal pathogen that targets over 500 plant species globally, including essential staple crops. Utilizing field image data for precise detection and quantification of CRS could greatly assist in the prompt identification and management of affected fields and thereby reduce yield losses. The objective of this work was to implement various machine learning algorithms to evaluate their ability to accurately detect and quantify CRS in red‐green‐blue images of sorghum plants exhibiting symptoms of infection. EfficientNet‐B3 and a fully convolutional network emerged as the top‐performing models for image classification and segmentation tasks, respectively. Among the classification models evaluated, EfficientNet‐B3 demonstrated superior performance, achieving an accuracy of 86.97%, a recall rate of 0.71, and an F1 score of 0.73. Of the segmentation models tested, FCN proved to be the most effective, exhibiting a validation accuracy of 97.76%, a recall rate of 0.68, and an F1 score of 0.66. As the size of the image patches increased, both models’ validation scores increased linearly, and their inference time decreased exponentially. This trend could be attributed to larger patches containing more information, improving model performance, and fewer patches reducing the computational load, thus decreasing inference time. The models, in addition to being immediately useful for breeders and growers of sorghum, advance the domain of automated plant phenotyping and may serve as a foundation for drone‐based or other automated field phenotyping efforts. Additionally, the models presented herein can be accessed through a web‐based application where users can easily analyze their own images.

Gonzalez, Emmanuel M.↗

Flying Blind, or Just Flying Under the Radar? The Underappreciated Power of De Novo Methods of Mass Spectrometric Peptide Identification

Mass spectrometry-based proteomics is a popular and powerful method for precise and highly multiplexed protein identification. The most common method of analyzing untargeted proteomics data is called database searching, where the database is simply a collection of protein sequences from the target organism, derived from genome sequencing. Experimental peptide tandem mass spectra are compared to simplified models of theoretical spectra calculated from the translated genomic sequences. However, in several interesting application areas, such as forensics, archaeology, venomics and others, a genome sequence may not be available, or the correct genome sequence to use is not known. In these cases, de novo peptide identification can play an important role. De novo peptide identification infers peptide sequence directly from the tandem mass spectrum without reference to a sequence database, usually using graph-based or machine learning algorithms. In this review, we provide a basic overview of de novo peptide identification methods and applications, briefly covering de novo algorithms and tools, and focusing in more depth on recent applications from venomics, metaproteomics, forensics, and characterization of antibody drugs.

proteomics, mass spectrometry, forensics, de novo ↗

Autonomous Aerosol and Plasma Co‐Jet Printing of Metallic Devices at Ambient Temperature

Abstract Additive manufacturing of metallic materials holds the potential to revolutionize the fabrication of functional devices unattainable via traditional methods. Despite recent advancements, printing metallic materials typically requires thermal processing at elevated temperatures to form dense structures with desired properties, which presents a major challenge for direct printing and integration with temperature‐sensitive materials. Herein, a unique co‐jet printing (CJP) method is reported integrating an aerosol jet and a non‐thermal, atmospheric pressure plasma jet to enable concurrent aerosol deposition of metal nanoparticle inks and in situ sintering at ambient temperature. A machine learning algorithm is integrated with the CJP to perform real‐time defect detection and autonomous correction, enhancing the yield of printed films with high electrical conductivity from 44% to 94%. Concurrent printing and sintering eliminate the need for post‐printing processing, reducing the overall manufacturing time by multiple folds depending on product size. CJP enables direct printing of functional devices on a variety of temperature‐sensitive materials including biological materials. Direct printing of hydration sensors on living plant leaves is demonstrated for long‐duration monitoring of hydration level in the plant. The versatile CJP method opens tremendous opportunities to harmoniously integrate abiotic and biotic materials for emerging applications in wearable/implantable devices and biohybrid systems.

Du, Yipu [Department of Aerospace and Mechanical E↗

New Directions for Thermoelectrics: A Roadmap from High‐Throughput Materials Discovery to Advanced Device Manufacturing

Thermoelectric materials, which can convert waste heat into electricity or act as solid‐state Peltier coolers, are emerging as key technologies to address global energy shortages and environmental sustainability. However, discovering materials with high thermoelectric conversion efficiency is a complex and slow process. The emerging field of high‐throughput material discovery demonstrates its potential to accelerate the development of new thermoelectric materials combining high efficiency and low cost. The synergistic integration of high‐throughput material processing and characterization techniques with machine learning algorithms can form an efficient closed‐loop process to generate and analyze broad datasets to discover new thermoelectric materials with unprecedented performances. Meanwhile, the recent development of advanced manufacturing methods provides exciting opportunities to realize scalable, low‐cost, and energy‐efficient fabrication of thermoelectric devices. This review provides an overview of recent advances in discovering thermoelectric materials using high‐throughput methods, including processing, characterization, and screening. Advanced manufacturing methods of thermoelectric devices are also introduced to realize the broad impacts of thermoelectric materials in power generation and solid‐state cooling. In the end, this article also discusses the future research prospects and directions.

Song, Kaidong↗

Importance of Depth and Artificial Structure as Predictors of Female Red Snapper Reproductive Parameters

Abstract The Red Snapper Lutjanus campechanus is a structure‐associated species occurring across a wide depth range in the northern Gulf of Mexico. We used the random forest machine learning algorithm to understand which habitat and individual fish characteristics could predict reproductive parameters of female Red Snapper. We evaluated fish captured from 2016 to 2018 on three artificial structure types with various structure heights at depths of 100 m or less. Overall, we found that depth and month were important predictors for most reproductive parameters, but the type of structure (artificial reefs, oil platforms, and rigs‐to‐reefs structures) was not important. Maturity was correctly classified in 88.9% of the cases when using the random forest ensemble model, with important predictors including FL, depth, structure height, and month of collection. Spawning seasonality (measured as gonadosomatic index [GSI]) was correctly classified in 59.5% of the cases when using histology reproductive phase, FL, month, and depth variables. Reproductively active or inactive females were correctly classified in 89.3% of the cases using GSI, month, FL, and depth, while females in the developing versus spawning capable phases were correctly classified in 82.2% of the cases using GSI, FL, month, and depth. Histological indicators that show potential spawning within a 36‐h period were correctly classified 61.5% of the time, with the best predictors being depth, FL, GSI, and month. Stepwise regression indicated that month was the only factor that significantly predicted contrasts in relative batch fecundity, with significantly greater values in August compared to all other months. Our findings suggest that female Red Snapper reproductive effort is not consistently or well predicted by artificial structure type or height but that a combination of fish FL, month, and depth can predict reproductive characteristics of female Red Snapper.

Brown‐Peterson, Nancy J.↗

Understanding and Leveraging the I/O Patterns of Emerging Machine Learning Analytics

The scientific community is currently experiencing unprecedented amounts of data generated by cutting-edge science facilities. Soon facilities will be producing up to 1 PB/s which will force scientist to use more autonomous techniques to learn from the data. The adoption of machine learning methods, like deep learning techniques, in large-scale workflows comes with a shift in the workflow’s computational and I/O patterns. These changes often include iterative processes and model architecture searches, in which datasets are analyzed multiple times in different formats with different model configurations in order to find accurate, reliable and efficient learning models. This shift in behavior brings changes in I/O patterns at the application level as well at the system level. These changes also bring new challenges for the HPC I/O teams, since these patterns contain more complex I/O workloads. In this paper we discuss the I/O patterns experienced by emerging analytical codes that rely on machine learning algorithms and highlight the challenges in designing efficient I/O transfers for such workflows. We comment on how to leverage the data access patterns in order to fetch in a more efficient way the required input data in the format and order given by the needs of the application and how to optimize the data path between collaborative processes. We will motivate our work and show performance gains with a study case of medical applications.

Gainaru, Ana↗

ReSpike: A Co-Design Framework for Evaluating SNNs on ReRAM-Based Neuromorphic Processors

With Moore’s law approaching its end, traditional von Neumann architectures are struggling to keep up with the exceeding performance and memory requirements of artificial intelligence and machine learning algorithms. Unconventional computing approaches such as neuromorphic computing that leverage spiking neural networks (SNNs) to perform computation are gaining traction and seek the paradigm shift necessary to sustain the increasing demands of modern applications. Novel memory technologies, such as resistive RAM (ReRAM), employ a crossbar architecture that possesses the inherent capability of efficiently computing vector-matrix multiplication—a dominant operation in SNNs. The prospect of naturally mapping SNNs to the crossbar structures provides a unique opportunity for achieving a high-performance, power-efficient neuromorphic system. In this work, we present ReSpike, which is a new framework, behavioral simulator, and architectural design based on ReRAM crossbar architectures, enabling modeling and co-design to achieve efficient execution of SNNs. We drive this co-design forward by quantifying the impact that ReRAM cell nonidealities have on the corresponding accuracy of an SNN application.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Real-time dynamics of the Schwinger model as an open quantum system with Neural Density Operators

Ab-initio simulations of multiple heavy quarks propagating in a Quark-Gluon Plasma are computationally difficult to perform due to the large dimension of the space of density matrices. This work develops machine learning algorithms to overcome this difficulty by approximating exact quantum states with neural network parametrisations, specifically Neural Density Operators. As a proof of principle demonstration in a QCD-like theory, the approach is applied to solve the Lindblad master equation in the 1 + 1d lattice Schwinger Model as an open quantum system. Neural Density Operators enable the study of in-medium dynamics on large lattice volumes, where multiple-string interactions and their effects on string-breaking and recombination phenomena can be studied. Thermal properties of the system at equilibrium can also be probed with these methods by variationally constructing the steady state of the Lindblad master equation. Scaling of this approach with system size is studied, and numerical demonstrations on up to 32 spatial lattice sites and with up to 3 interacting strings are performed.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nano-enhanced solid-state hydrogen storage: Balancing discovery and pragmatism for future energy solutions

Nanomaterials have revolutionized the battery industry by enhancing energy storage capacities and charging speeds, and their application in hydrogen (H 2 ) storage likewise holds strong potential, though with distinct challenges and mechanisms. H 2 is a crucial future zero-carbon energy vector given its high gravimetric energy density, which far exceeds that of liquid hydrocarbons. However, its low volumetric energy density in gaseous form currently requires storage under high pressure or at low temperature. This review critically examines the current and prospective landscapes of solid-state H 2 storage technologies, with a focus on pragmatic integration of advanced materials such as metal-organic frameworks (MOFs), magnesium-based hybrids, and novel sorbents into future energy networks. These materials, enhanced by nanotechnology, could significantly improve the efficiency and capacity of H 2 storage systems by optimizing H 2 adsorption at the nanoscale and improving the kinetics of H 2 uptake and release. We discuss various H 2 storage mechanisms—physisorption, chemisorption, and the Kubas interaction—analyzing their impact on the energy efficiency and scalability of storage solutions. The review also addresses the potential of “smart MOFs”, single-atom catalyst-doped metal hydrides, MXenes and entropy-driven alloys to enhance the performance and broaden the application range of H 2 storage systems, stressing the need for innovative materials and system integration to satisfy future energy demands. High-throughput screening, combined with machine learning algorithms, is noted as a promising approach to identify patterns and predict the behavior of novel materials under various conditions, significantly reducing the time and cost associated with experimental trials. In closing, we discuss the increasing involvement of various companies in solid-state H 2 storage, particularly in prototype vehicles, from a techno-economic perspective. In conclusion, this forward-looking perspective underscores the necessity for ongoing material innovation and system optimization to meet the stringent energy demands and ambitious sustainability targets increasingly in demand.

25 ENERGY STORAGE↗

Relevant biochar characteristics influencing compressive strength of biochar-cement mortars

To counteract the contribution of CO 2 emissions by cement production and utilization, biochar is being harnessed as a carbon-negative additive in concrete. Increasing the cement replacement and biochar dosage will increase the carbon offset, but there is large variability in methods being used and many researchers report strength decreases at cement replacements beyond 5%. This work presents a reliable method to replace 10% of the cement mass with a vast selection of biochars without decreasing ultimate compressive strength, and in many cases significantly improving it. By carefully quantifying the physical and chemical properties of each biochar used, machine learning algorithms were used to elucidate the three most influential biochar characteristics that control mortar strength: initial saturation percentage, oxygen-to-carbon ratio, and soluble silicon. These results provide additional research avenues for utilizing several potential biomass waste streams to increase the biochar dosage in cement mixes without decreasing mechanical properties.

97 MATHEMATICS AND COMPUTING↗

A systematic analysis of phase stability in refractory high entropy alloys utilizing linear and non-linear cluster expansion models

We report that obtaining surrogate models for various types of materials has long been a goal of computational material scientists and physicists, as they allow the calculation of thermodynamic quantities and properties of interest far more efficiently than DFT. Many surrogate models, including interatomic potentials generated by machine learning algorthims and cluster expansions, rely on incorporating many-body interactions such as three-body or four-body interactions. When utilizing machine learning algorithms for interatomic potentials, it is common to start with two-body interactions and systematically include three- and four-body clusters to improve the fit, stopping when including the additional clusters does not measurably change the predictive power of the model. In cluster expansions for FCC and related lattices, it is often enough to include mostly two-body interactions and a handful of three-body interactions. However, it remains to be seen whether those will be enough in the case of lattices with a lower packing fraction, such as BCC.

36 MATERIALS SCIENCE↗

Deep reinforcement learning for class imbalance fault diagnosis of equipment in nuclear power plants

In equipment fault diagnosis in nuclear power plants, there may be far more samples in one class (e.g., a health state) than in another class (e.g., a fault state). The distribution of data in each class is highly skewed. Most machine learning algorithms are suitable for balanced training datasets. When faced with imbalanced samples, these algorithms tend to provide good identification for the majority classes and bias for the minority classes. However, the misclassification of minority classes can lead to high costs. To address the above problem, this paper develops a deep reinforcement learning-based diagnosis method that models fault diagnosis as a sequential decision-making process. At each time step, the agent receives the state of the environment represented by the training samples and then takes a diagnosis action guided by a policy. If the action is correct/incorrect, the agent receives a positive/negative reward. The reward for minority classes is higher than that for majority classes. The agent’s goal is to obtain as many cumulative rewards as possible in the process, i.e., to identify the sample as correctly as possible. Six demonstration scenarios are constructed, depending on the selected fault datasets and the designed model structures. Experiments show that the proposed method achieves a higher weighted-averaged F1 score than the classical supervised learning method in most cases of class imbalance. Finally, the proposed method has potential applications in the field of class imbalance fault diagnosis of equipment in nuclear power plants.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Hydrogen production from lignocellulosic hydrolysate in an up-scaled microbial electrolysis cell with stacked bio-electrodes

Hydrogen production from renewable resources via microbial electrolysis cells (MECs) is a promising approach for sustainable energy production. Yet high hydrogen yield from real feedstocks has not been demonstrated in up-scaled MECs. In this study, a 10-L single chamber MEC with a high electrode surface area to volume ratio (66 m 2 /m 3 ) was constructed and electroactive cathodic biofilms were enriched for hydrogen evolution reaction. A high hydrogen yield of 91% was achieved using lignocellulosic hydrolysate with a hydrogen production rate of 0.71 L/L/D at an organic loading rate of 0.4 g/D. The anodic and cathodic microbial communities, with Enterococcus spp. as the known electroactive bacteria, were capable of achieving current densities of 13.7 A/m 2 and 16.5 A/m 2 , respectively. Finally, a machine learning algorithm was used to investigate the correlation between community data and electrochemical performance, and the critical genera on determining current density were identified.

54 ENVIRONMENTAL SCIENCES↗

Characterization data of an (AlFeNiTiVZr) 1-x Cr x multi-principal element alloy continuous composition spread library

The data provided in this article is related to the research article entitled “Phase stabilization and oxidation of a continuous composition spread multi-principal element (AlFeNiTiVZr) 1-x Cr x alloy”. This data article describes the high-throughput synthesis and characterization processes of an (AlFeNiTiVZr) 1-x Cr x alloy system. Continuous composition spread (CCS) thin-film libraries were synthesized by co-depositing an AlFeNiTiVZr metal alloy target and Cr target via magnetron sputtering. Post-processing was performed on the sample libraries with a vacuum anneal at 873 K and an air anneal at 873 K. Compositional data was determined via WDS in order to verify parameters provided by an in-house sputter model. Crystallographic data was captured via synchrotron diffraction and diffractograms were compared as a function of the change in Cr concentration. These measurements were taken in order to observe phase behavior after oxidation throughout the composition library. Furthermore, vibrational spectrographic data is provided of the oxidized library to show surface speciation along the composition gradient of the alloy system. The structural and oxidative behavior of the (AlFeNiTiVZr) 1-x Cr x alloy can be analysed using the data provided in this article. Additionally, this characterization dataset can be utilized in machine learning algorithms for determining important features and parameters for future hypothesis generation of functional multi-principal element alloys (MPEAs).

36 MATERIALS SCIENCE↗

Adaptive Algebraic Derivative Estimation for Battery Electric Buses Energy Consumption Forecasting

The limited service life of onboard batteries for EVs is a challenge, underscoring the need for real-time battery usage prediction. This paper proposes an adaptive Algebraic Derivative Estimation (ADE) approach for forecasting the energy consumption of battery electric buses. By dynamically adjusting the sliding window length, the adaptive ADE retains the fixed-length ADE’s key advantage—namely, operating online without reliance on extensive historical datasets—while substantially bolstering forecast accuracy by actively trading estimation bias off estimation variance. Comparative experiments against both the conventional ADE with a fixed length and a representative machine learning algorithm, XGBoost, were conducted, with performance evaluated via root mean square error, mean absolute error, and the coefficient of determination. The results demonstrate that the proposed approach significantly outperforms baseline methods.

Cui, Tianyang [The University of Texas at Dallas]↗

Deep learning uncertainty quantification for clinical text classification

Machine learning algorithms are expected to work side-by-side with humans in decision-making pipelines. Thus, the ability of classifiers to make reliable decisions is of paramount importance. Deep neural networks (DNNs) represent the state-of-the-art models to address real-world classification. Although the strength of activation in DNNs is often correlated with the network’s confidence, in-depth analyses are needed to establish whether they are well calibrated. In this paper, we demonstrate the use of DNN-based classification tools to benefit cancer registries by automating information extraction of disease at diagnosis and at surgery from electronic text pathology reports from the US National Cancer Institute (NCI) Surveillance, Epidemiology, and End Results (SEER) population-based cancer registries. In particular, we introduce multiple methods for selective classification to achieve a target level of accuracy on multiple classification tasks while minimizing the rejection amount—that is, the number of electronic pathology reports for which the model’s predictions are unreliable. We evaluate the proposed methods by comparing our approach with the current in-house deep learning-based abstaining classifier. Overall, all the proposed selective classification methods effectively allow for achieving the targeted level of accuracy or higher in a trade-off analysis aimed to minimize the rejection rate. On in-distribution validation and holdout test data, with all the proposed methods, we achieve on all tasks the required target level of accuracy with a lower rejection rate than the deep abstaining classifier (DAC). Interpreting the results for the out-of-distribution test data is more complex; nevertheless, in this case as well, the rejection rate from the best among the proposed methods achieving 97% accuracy or higher is lower than the rejection rate based on the DAC. We show that although both approaches can flag those samples that should be manually reviewed and labeled by human annotators, the newly proposed methods retain a larger fraction and do so without retraining—thus offering a reduced computational cost compared with the in-house deep learning-based abstaining classifier.

59 BASIC BIOLOGICAL SCIENCES↗

Nonlinear sparse Bayesian learning for physics-based models

This paper addresses the issue of overfitting while calibrating unknown parameters of over-parameterized physics-based models with noisy and incomplete observations. Here, a semi-analytical Bayesian framework of nonlinear sparse Bayesian learning (NSBL) is proposed to identify sparsity among model parameters during Bayesian inversion. NSBL offers significant advantages over machine learning algorithm of sparse Bayesian learning (SBL) for physics-based models, such as 1) the likelihood function or the posterior parameter distribution is not required to be Gaussian, and 2) prior parameter knowledge is incorporated into sparse learning (i.e. not all parameters are treated as questionable). NSBL employs the concept of automatic relevance determination (ARD) to facilitate sparsity among questionable parameters through parameterized prior distributions. The analytical tractability of NSBL is enabled by employing Gaussian ARD priors and by building a Gaussian mixture-model approximation of the posterior parameter distribution that excludes the contribution of ARD priors. Subsequently, type-II maximum likelihood is executed using Newton's method whereby the evidence and its gradient and Hessian information are computed in a semi-analytical fashion. We show numerically and analytically that SBL is a special case of NSBL for linear regression models. Subsequently, a linear regression example involving multimodality in both parameter posterior pdf and model evidence is considered to demonstrate the performance of NSBL in cases where SBL is inapplicable. Next, NSBL is applied to identify sparsity among the damping coefficients of a mass-spring-damper model of a shear building frame. These numerical studies demonstrate the robustness and efficiency of NSBL in alleviating overfitting during Bayesian inversion of nonlinear physics-based models.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗