Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Randomized methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Factors influencing mode choice of adults with travel-limiting disability

Introduction: Despite the plethora of research devoted to analyzing the impact of disability on travel behavior, not enough studies have investigated the varying impact of social and environmental factors on the mode choice of people with disabilities that restrict their ability to use transportation modes efficiently. This research gap can be addressed by investigating the factors influencing the mode choice behavior of people with travel-limiting disabilities, which can inform the development of accessible and sustainable transportation systems. Additionally, such studies can provide insights into the social and economic barriers faced by this population group, which can help policymakers to promote social inclusion and equity. Method: This study utilized a Random Parameters Logit model to identify the individual, trip, and environmental factors that influence mode selection among people with travel-limiting disabilities. Here, using the 2017 National Household Travel Survey data for New York State, which included information on respondents with travel-limiting disabilities, the analysis focused on a sample of 8,016 people. In addition, climate data from the National Oceanic and Atmospheric Administration were integrated as additional explanatory variables in the modeling process. Results: The results revealed that people with disabilities may be inclined to travel longer distances walking in the absence of suitable accommodation facilities for other transportation modes. Furthermore, people were less inclined to walk during summer and winter, indicating a need to consider weather conditions as a significant determinant of mode choice. Moreover, low-income people with disabilities were more likely to rely on public transport or walking. Conclusion: Based on this study’s findings, transportation agencies could design infrastructure and plan for future expansions that is more inclusive and accessible, thus catering to the mobility needs of people with travel-limiting disabilities.

99 GENERAL AND MISCELLANEOUS↗

Modeling Nanoconfinement Effects Using Active Learning

Predicting the spatial configuration of gas in nanopores of is relevant in applications such as fluid flow forecasting and hydrocarbon reserves estimation. For example, shale reservoirs have suffered from computationally intractable multiscale problems, since fluid properties such as viscosity, density, and adsorption must be calculated by using expensive molecular dynamics (MD) simulations within each nanopore, whereas flow through these connected nanopores must be simulated at the micrometer scale. We utilize machine learning techniques to quickly and accurately model nanoscale confinement effects as an important step toward bridging the nano and micro scales. Our workflow is based on building and training physics-based deep-neural-networks models by learning from a database of MD calculations. The model accounts for the adsorption phenomenon by predicting the statistical distribution of gas inside nanopores. Because large databases of MD calculations are expensive to create, we investigate active learning (AL) as a data set construction strategy. In this workflow, new data are selected based on the model uncertainty via the query-by-committee approach. We show that our workflow obtains accurate models that generalize to real scanning electron microscopy geometries with 1/10th of the number of MD calculations required vs random data set generation. Our method enables the possibility of modeling nanoconfinement effects at the mesoscale, where complex connected sets of nanopores affect flow.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Functionalized Graphene via a One-Pot Reaction Enabling Exact Pore Sizes, Modifiable Pore Functionalization, and Precision Doping

Functionalizing graphene with exact pore size, specific functional groups, and precision doping poses many significant challenges. Current methods lack precision and produce random pore sizes, sites of attachment, and amounts of dopant, leading to compromised structural integrity and affecting graphene’s applications. In this work, we report a strategy for the synthesis of functionalized graphitic materials with modifiable nanometer-sized pores via a Pictet–Spengler polymerization reaction. This one-pot, four-step synthesis uses concepts based on covalent organic frameworks (COFs) synthesis to produce crystalline two-dimensional materials that were confirmed by PXRD, TEM measurements, and DFT studies. These new materials are structurally analogous to doped graphene and graphene oxide (GO) but, unlike GO, maintain their semiconductive properties when fully functionalized.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Consistent Representation of Cloud Overlap and Cloud Subgrid Vertical Heterogeneity

Many global climate models underestimate the cloud cover and overestimate the cloud albedo, especially for low-level clouds. We determine how a correct representation of the vertical structure of clouds can fix part of this bias. We use the 1D McICA framework and focus on low-level clouds. Using Large Eddy Simulations results as reference, we propose a method based on exponential-random overlap that represents the cloud overlap between layers and the subgrid cloud properties over several vertical scales, with a single value of the overlap parameter. Starting from a coarse vertical grid, representative of atmospheric models, this algorithm is used to generate the vertical profile of the cloud fraction with a finer vertical resolution, or to generate it on the coarse grid but with subgrid heterogeneity and cloud overlap that ensures a correct cloud cover. Doing so we find decorrelation lengths are dependent on the vertical resolution, except if the vertical subgrid heterogeneity and interlayer overlap are taken into account coherently. We confirm that the frequently used maximum-random overlap leads to a significant error by underestimating the low-level cloud cover with a relative error of about 50%, that can lead to an error of SW cloud albedo as big as 70%. Not taking into account the subgrid vertical heterogeneity of clouds can cause a relative error of 20% in brightness, assuming the cloud cover is correct.

54 ENVIRONMENTAL SCIENCES↗

Simultaneous quantification of uranium( VI ), samarium, nitric acid, and temperature with combined ensemble learning, laser fluorescence, and Raman scattering for real-time monitoring

In this work, laser-induced fluorescence spectroscopy (LIFS), Raman spectroscopy, and a stacked regression ensemble was developed for near real-time quantification of uranium(VI) (1–100 μg mL –1 ), samarium (0–200 μg mL –1 ) and nitric acid (0.1–4 M) with varying temperature (20 °C–45 °C). LIFS applications range from fundamental lab-scale studies to real-time process monitoring at industrial levels, such as nuclear reprocessing applications, provided the phenomena affecting the fluorescence spectrum are accounted for (e.g., absorption, quenching, complexation). Multiple chemometric models were examined and compared to a more traditional multivariate regression approach called partial least squares (PLS). Results obtained on synthetic samples selected using D-optimal experimental design indicated that a stacked regression method, which included ridge regression, random forest, PLS, and an eXtreme gradient boost algorithm, successfully measured uranium(VI) concentrations directly in nitric acid without measuring luminescence lifetimes or standard addition. The top model resulted in percent root-mean-square error of prediction values of 5.2, 1.9, 3.0, and 2.3% for U(VI), Sm 3+ , HNO 3 , and temperature, respectively. The approach may be useful for quantifying fluorescent fission products (e.g., Sm 3+ ) to provide information on burnup of irradiated nuclear fuel. This novel framework reinforces the applicability of LIFS for real-time applications in nuclear fuel cycle applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Controllable Reset Behavior in Domain Wall–Magnetic Tunnel Junction Artificial Neurons for Task-Adaptable Computation

Neuromorphic computing with spintronic devices has been of interest due to the limitations of CMOS-driven von Neumann computing. Domain wall–magnetic tunnel junction (DW-MTJ) devices have been shown to be able to intrinsically capture biological neuron behavior. Edgy-relaxed behavior, where a frequently firing neuron experiences a lower action potential threshold, may provide additional artificial neuronal functionality when executing repeated tasks. In this letter, we demonstrate that this behavior can be implemented in DW-MTJ artificial neurons via three alternative mechanisms: shape anisotropy, magnetic field, and current-driven soft reset. Using micromagnetics and analytical device modeling to classify the Optdigits handwritten digit dataset, we show that edgy-relaxed behavior improves both classification accuracy and classification rate for ordered datasets while sacrificing little to no accuracy for a randomized dataset. This letter establishes methods by which artificial spintronic neurons can be flexibly adapted to datasets.

42 ENGINEERING↗

DT-HYDRO

The software solves the time dependent, one dimensional (1D) coupled mass and momentum balance equations governing the elastic flow of water through the penstock, turbine and draft tube in a hydroelectric facility using a high order finite volume based method. The numerical method is based on the Kurganov-Tadmor central method paired with the Monotonic Upstream-centered Scheme for Conservation Laws (MUSCL). This solution method accurately resolves the fast transient behavior of the flow, including water hammer. Additionally, the software estimates the full 3D flow field within the turbine chamber in real time, a feat that is made possible by leveraging pre-computed CFD results by utilizing a reduced order modeling method based on an efficient randomized singular value decomposition (SVD) driven proper orthogonal decomposition (POD) with POD-mode weight regression. The reduced order model of the 3D flow is directly coupled to the 1D elastic flow model so the entire flow field through the penstock and turbine system is resolved quickly and with high fidelity.

Gurecky, William [Oak Ridge National Laboratory (O↗

Data Science and Machine Learning for Genome Security

This report describes research conducted to use data science and machine learning methods to distinguish targeted genome editing versus natural mutation and sequencer machine noise. Genome editing capabilities have been around for more than 20 years, and the efficiencies of these techniques has improved dramatically in the last 5+ years, notably with the rise of CRISPR-Cas technology. Whether or not a specific genome has been the target of an edit is concern for U.S. national security. The research detailed in this report provides first steps to address this concern. A large amount of data is necessary in our research, thus we invested considerable time collecting and processing it. We use an ensemble of decision tree and deep neural network machine learning methods as well as anomaly detection to detect genome edits given either whole exome or genome DNA reads. The edit detection results we obtained with our algorithms tested against samples held out during training of our methods are significantly better than random guessing, achieving high F1 and recall scores as well as with precision overall.

59 BASIC BIOLOGICAL SCIENCES↗

Monte Carlo Transport: Computational Physics Summer Workshop [Slides]

In general, Monte Carlo methods simulate large numbers of random trials in order to observe numerical behavior of systems described by probabilistic behavior. In radiation transport, pseudo-random number generators are used to randomly sample individual particle lives. Information about the particles are tallied.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Efficient QAOA Optimization using Directed Restarts and Graph Lookup

Variational Quantum Algorithms (VQA) aim to enhance the capabilities of Noisy Intermediate-Scale Quantum (NISQ) devices. These algorithms utilize parameterized circuits and classical optimizers to iteratively execute circuits with varying parameters. However, VQA faces computational overheads due to repeated iterations and random restarts. Prior work suggests using basic sub-graphs to transfer parameters for the input graph, reducing optimizer overheads but limiting applicability to structured regular graphs. In real-world applications, random irregular graphs are common, and existing methods are not scalable or practical for such graphs. This paper presents a framework that aims to improve random irregular graphs in VQA. The framework uses graph similarity and important features like total edge counts, average edge counts, and variance. It follows an iterative process to choose basis sub-graphs from a small database and adjust parameters accordingly. Classical optimizers then utilize these parameters to determine when to restart and perform gradient descent. This approach increases the chances of reaching global maximum points.

Wang, Meng↗

Improved Subseasonal Forecasting of Extreme Polar Vortices Using Machine Learning

Our research was focused on forecasting the position and shape of the winter stratospheric polar vortex at a subseasonal timescale of 15 days in advance. To achieve this, we employed both statistical and neural network machine learning techniques. The analysis was performed on 42 winter seasons of reanalysis data provided by NASA giving us a total of 6,342 days of data. The state of the polar vortex for determined by using geometric moments to calculate the centroid latitude and the aspect ratio of an ellipse fit onto the vortex. Timeseries for thirty additional precursors were calculated to help improve the predictive capabilities of the algorithm. Feature importance of these precursors was performed using random forest to measure the predictive importance and the ideal number of precursors. Then, using the precursors identified as important, various statistical methods were tested for predictive accuracy with random forest and nearest neighbor performing the best. An echo state network, a type of recurrent neural network that features sparsely connected hidden layer and a reduced number of trainable parameters that allows for rapid training and testing, was also implemented for the forecasting problem. Hyperparameter tuning was performed for each methods using a subset of the training data. The algorithms were trained and tuned on the first 41 years of data, then tested for accuracy on the final year. In general, the centroid latitude of the polar vortex proved easier to predict than the aspect ratio across all algorithms. Random forest outperformed other statistical forecasting algorithms overall but struggled to predict extreme values. Forecasting from echo state network suggested a strong predictive capability past 15 days, but further work is required to fully realize the potential of recurrent neural network approaches.

54 ENVIRONMENTAL SCIENCES↗

Effective Field Theory of Random Quantum Circuits

Quantum circuits have been widely used as a platform to simulate generic quantum many-body systems. In particular, random quantum circuits provide a means to probe universal features of many-body quantum chaos and ergodicity. Some such features have already been experimentally demonstrated in noisy intermediate-scale quantum (NISQ) devices. On the theory side, properties of random quantum circuits have been studied on a case-by-case basis and for certain specific systems, and a hallmark of quantum chaos—universal Wigner–Dyson level statistics—has been derived. This work develops an effective field theory for a large class of random quantum circuits. The theory has the form of a replica sigma model and is similar to the low-energy approach to diffusion in disordered systems. The method is used to explicitly derive the universal random matrix behavior of a large family of random circuits. In particular, we rederive the Wigner–Dyson spectral statistics of the brickwork circuit model by Chan, De Luca, and Chalker [Phys. Rev. X 8, 041019 (2018)] and show within the same calculation that its various permutations and higher-dimensional generalizations preserve the universal level statistics. Finally, we use the replica sigma model framework to rederive the Weingarten calculus, which is a method of evaluating integrals of polynomials of matrix elements with respect to the Haar measure over compact groups and has many applications in the study of quantum circuits. The effective field theory derived here provides both a method to quantitatively characterize the quantum dynamics of random Floquet systems (e.g., calculating operator and entanglement spreading) and a path to understanding the general fundamental mechanism behind quantum chaos and thermalization in these systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Posiform planting: generating QUBO instances for benchmarking

We are interested in benchmarking both quantum annealing and classical algorithms for minimizing quadratic unconstrained binary optimization (QUBO) problems. Such problems are NP-hard in general, implying that the exact minima of randomly generated instances are hard to find and thus typically unknown. While brute forcing smaller instances is possible, such instances are typically not interesting due to being too easy for both quantum and classical algorithms. In this contribution, we propose a novel method, called posiform planting , for generating random QUBO instances of arbitrary size with known optimal solutions, and use those instances to benchmark the sampling quality of four D-Wave quantum annealers utilizing different interconnection structures (Chimera, Pegasus, and Zephyr hardware graphs) and the simulated annealing algorithm. Posiform planting differs from many existing methods in two key ways. It ensures the uniqueness of the planted optimal solution, thus avoiding groundstate degeneracy, and it enables the generation of QUBOs that are tailored to a given hardware connectivity structure, provided that the connectivity is not too sparse. Posiform planted QUBOs are a type of 2-SAT boolean satisfiability combinatorial optimization problems. Our experiments demonstrate the capability of the D-Wave quantum annealers to sample the optimal planted solution of combinatorial optimization problems with up to 5, 627 qubits.

97 MATHEMATICS AND COMPUTING↗

A spatially regularized detector for emergent/re-emergent disease outbreaks

Early detection of outbreaks caused by emergent pathogens, using epidemiological surveillance data i.e., daily case counts, is difficult. This is because the data tend to be noisy during the early epoch of the outbreak. In contrast, the spread-rate of the disease tends to be well-behaved, as it depends only on the mixing patterns of the population and the characteristics of the pathogen, neither of which behave erratically in space-time. In this report, we explore whether the spread-rate can be used for epidemiological surveillance, conditional on case count data. Estimating the spread-rate from case count data allows us to exploit exogenous information, e.g., incubation period distributions etc., which can considerably smooth out any erratic temporal behavior. Further, epidemiological dynamics are spatially correlated, and if case counts are available for multiple areal units e.g., counties, these correlations could potentially be used to suppress noise in the early epoch data. These exogenous information and structure are not exploited by conventional syndromic surveillance detectors to extract a well-behaved latent variable for monitoring purposes. The technical challenge lies in the estimation of the spread-rate field jointly over a collection of areal units; further, the spread-rate varies over time. We develop a method based on mean-field variational inference to approximately estimate the spread-rate field, using a Gaussian Random Field Model for spatial regularization. The method is tested on the estimation of spread-rate in the thirty-three counties of New Mexico and detect the arrival of the Fall 2020 COVID-19 wave in September 2020. We find that the method is scalable, but underestimates the uncertainty in the estimated spread-rate field. We detect the arrival of the Fall 2020 wave a week ahead of conventional syndromic surveillance algorithms, but our simplistic detection algorithm, based on simple anomaly detection, suffers from a high false positive rate, similar to conventional detectors.

59 BASIC BIOLOGICAL SCIENCES↗

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗

Unifying and benchmarking state-of-the-art quantum error mitigation techniques

Error mitigation is an essential component of achieving a practical quantum advantage in the near term, and a number of different approaches have been proposed. In this work, we recognize that many state-of-the-art error mitigation methods share a common feature: they are data-driven, employing classical data obtained from runs of different quantum circuits. For example, Zero-noise extrapolation (ZNE) uses variable noise data and Clifford-data regression (CDR) uses data from near-Clifford circuits. We show that Virtual Distillation (VD) can be viewed in a similar manner by considering classical data produced from different numbers of state preparations. Observing this fact allows us to unify these three methods under a general data-driven error mitigation framework that we call UNIfied Technique for Error mitigation with Data (UNITED). In certain situations, we find that our UNITED method can outperform the individual methods (i.e., the whole is better than the individual parts). Specifically, we employ a realistic noise model obtained from a trapped ion quantum computer to benchmark UNITED, as well as other state-of-the-art methods, in mitigating observables produced from random quantum circuits and the Quantum Alternating Operator Ansatz (QAOA) applied to Max-Cut problems with various numbers of qubits, circuit depths and total numbers of shots. We find that the performance of different techniques depends strongly on shot budgets, with more powerful methods requiring more shots for optimal performance. For our largest considered shot budget (10 10 ), we find that UNITED gives the most accurate mitigation. Hence, our work represents a benchmarking of current error mitigation methods and provides a guide for the regimes when certain methods are most useful.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Informing Plant Asset Reliability and Availability Through AI-Driven Analysis of Operator Logs

The availability and reliability of nuclear power plant (NPP) structures, systems, and components (SSCs) are critical parameters for NPP safety. Tracking these parameters is necessary but costly and labor-intensive, requiring the collection and evaluation of SSC event data such as shutdowns, startups, and failures. To show how these events are needed for the parameters an example is given: one measure of reliability is based on the number of equipment failure events and the number of run hours (i.e., the time from a startup event to a shutdown event). Here, this work investigates using artificial intelligence (AI) to mine NPP operator log entry texts for SSC event data. Four AI approaches were explored for identifying these events, including natural language processing (NLP) methods, generative AI, generative AI combined with NLP, and topic modeling. A key challenge addressed with all four approaches is the brevity of operator log entries. Among these four a neural network–based NLP method was shown to be the most promising for this application, achieving F1 scores of 86.0% for shutdowns, 92.2% for startups, and 80.4% for failures on a subject-matter-expert-curated dataset from NPP operator logs, compared to a baseline of 66.6% for a random classifier. This shows that NLP methods can perform better than generative AI. Additionally, the NLP methods combined with generative AI were shown to perform better than generative AI alone. Generative AI was most successful at providing the background information for the NLP methods to use. This work demonstrates the potential to use AI to automate parameter collection from NPP operator log entries and other records.

97 - MATHEMATICS AND COMPUTING↗