Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Randomized methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

On unifying randomized methods for inverse problems

This work unifies the analysis of various randomized methods for solving linear and nonlinear inverse problems with Gaussian priors by framing the problem in a stochastic optimization setting. By doing so, we show that many randomized methods are variants of a sample average approximation (SAA). More importantly, we are able to prove a single theoretical result that guarantees the asymptotic convergence for a variety of randomized methods. Additionally, viewing randomized methods as an SAA enables us to prove, for the first time, a single non-asymptotic error result that holds for randomized methods under consideration. Another important consequence of our unified framework is that it allows us to discover new randomization methods. Here, we present various numerical results for linear, nonlinear, algebraic, and PDE-constrained inverse problems that verify the theoretical convergence results and provide a discussion on the apparently different convergence rates and the behavior for various randomized methods.

42 ENGINEERING↗

Negative fluxes and cell-miss errors in the random ray method

The random ray method is a recently developed stochastic method for solving neutral particle transport problems based on the method of characteristics. Perhaps surprisingly for a characteristics-based method using flat sources, we note that the random ray method can produce negative fluxes which may be numerically troublesome in several situations. These occur most severely in fixed source problems where the source is in a region with a small cross section. Additionally, we briefly discuss another source of bias which can occur in similar situations, namely a ray missing a mesh with a strong source and small cross section, resulting in the entirety of the source being unphysically deposited locally. This paper describes the mechanism by which negative fluxes may occur and several different methods to mitigate their effects. These fixes are tested on an eigenvalue problem, a ‘fusion-like’ shielding problem, and a shielding problem featuring an adjoint calculation. Even when extremely coarse random ray quadratures are used such that 20%–30% of cells are missed during a given iteration, use of the preferred fix technique ensures local flux tally errors remain trivial (below 1%). The preferred fix is now the default option in SCONE and OpenMC.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Random Ray Method Versus Multigroup Monte Carlo: The Method of Characteristics in OpenMC and SCONE

The Random Ray Method (TRRM) is a recently developed approach to solving neutral particle transport problems based on the Method of Characteristics. While the method previously has been implemented only in closed-source or limited-functionality codes, this work describes its implementation in two open-source Monte Carlo codes: OpenMC and SCONE. The random ray implementations required small modifications to the existing Multigroup Monte Carlo (MGMC) solvers, offering a rare venue for redundant, fine-grained, "apples-to-apples" speed and accuracy comparisons between transport methods. To this end, TRRM and MGMC solvers are evaluated against each other using each code's native capabilities on reactor eigenvalue problems with different degrees of energy discretization. On the C5G7 benchmark (featuring only seven energy groups), TRRM achieves a maximum pin power error comparable to or lower than that of MGMC for a given run time. On a problem with 69 energy groups, MGMC is found to scale more efficiently, obtaining a lower pin power error for a given run time. However, the defining difference between the two transport methods is found to be their vastly different uncertainty distributions. Specifically, TRRM is found to maintain similar levels of accuracy and uncertainty throughout the simulation domain whereas MGMC can exhibit orders-of-magnitude greater errors in areas of the problem that feature low neutron flux. For instance, TRRM provided an up to 373 times speed advantage compared with MGMC for computing the flux in low-flux regions in the moderator surrounding the C5G7 core.

42 ENGINEERING↗

A mathematical assessment of the isolation random forest method for anomaly detection in big data

We present the mathematical analysis of the Isolation Random Forest Method (IRF Method) for anomaly detection, proposed by Liu F.T., Ting K.M. and Zhou Z. H. in their seminal work as a heuristic method for anomaly detection in Big Data. We prove that the IRF space can be endowed with a probability induced by the Isolation Tree algorithm (iTree). In this setting, the convergence of the IRF method is proved, using the Law of Large Numbers. Here, a couple of counterexamples are presented to show that the method is inconclusive and no certificate of quality can be given, when using it as a means to detect anomalies. Hence, an alternative version of the method is proposed whose mathematical foundation is fully justified. Furthermore, a criterion for choosing the number of sampled trees needed to guarantee confidence intervals of the numerical results is presented. Finally, numerical experiments are presented to compare the performance of the classic method with the proposed one.

97 MATHEMATICS AND COMPUTING↗

Automated ICRF heating surrogate modeling via machine learning

This work introduces automated machine learning workflows that address critical bottlenecks in surrogate model development for Ion Cyclotron Range of Frequencies (ICRF) heating applications. The automated framework includes data analysis tools that transform raw datasets into actionable insights in seconds, replacing weeks of manual exploratory effort and ensuring consistent, reproducible dataset characterization. By integrating advanced hyperparameter optimization (HPO) methods including Bayesian optimization via BoTorch and Tree-structured Parzen Estimators (TPE), the framework significantly reduces model development time from weeks to hours, decreasing computational cost and required expertise, while enabling high-accuracy surrogate models. Compared to traditional hyperparameter scanning (HPS) techniques such as methodical, randomized, and grid searches, HPO methods achieve superior convergence and predictive performance, even when compared to already well-tuned reference models. On NSTX High Harmonic Fast Wave (HHFW) heating datasets, both Random Forest Regressor (RFR) and neural network surrogates demonstrate improved accuracy, achieving R 2 values beyond 0.97 and 0.98, respectively. The results show that while HPO gains are modest for robust architectures like RFR, they become essential for more sensitive models such as neural networks, highlighting the trade-offs across optimization strategies. Through automated workflows that eliminate manual hyperparameter tuning and require minimal ML expertise, this work enables widespread adoption of high-fidelity surrogate models across the fusion community for real-time plasma control, uncertainty quantification, rapid experimental scenario development, and integrated system optimization.

Sanchez-Villar, Alvaro [Princeton Plasma Physics L↗

Systems and methods for randomized energy draw or supply requests

The present disclosure can provide a distributed and anonymous approach to demand response of an electricity system. The approach can conceptualize energy consumption and production of distributed-energy resources (DERs) via discrete energy packets that are coordinated by a cyber computing entity that grants or denies energy packet requests from the DERs. The approach leverages a condition of a DER, which is particularly useful for (1) thermostatically-controlled loads, (2) non-thermostatic conditionally-controlled loads, and (3) bi-directional distributed energy storage systems, among others. In a first aspect of the present approach, each DER independently requests the authority to switch on for a fixed amount of time (i.e., packet duration). The coordinator determines whether to grant or deny each request based electric grid and/or energy or power market conditions. In a second aspect, bi-directional DERs, such as distributed-energy storage systems (DESSs) are further able to request to supply energy to the grid.

Frolik, Jeff↗

Gradient Coding With Iterative Block Leverage Score Sampling

Gradient coding is a method for mitigating straggling servers in a centralized computing network that uses erasure-coding techniques to distributively carry out first-order optimization methods. Randomized numerical linear algebra uses randomization to develop improved algorithms for large-scale linear algebra computations. In this study, we propose a method for distributed optimization that combines gradient coding and randomized numerical linear algebra. The proposed method uses a randomized ℓ 2 -subspace embedding and a gradient coding technique to distribute blocks of data to the computational nodes of a centralized network, and at each iteration the central server only requires a small number of computations to obtain the steepest descent update. The novelty of our approach is that the data is replicated according to importance scores, called block leverage scores, in contrast to most gradient coding approaches that uniformly replicate the data blocks. Furthermore, we do not require a decoding step at each iteration, avoiding a bottleneck in previous gradient coding schemes. We show that our approach results in a valid ℓ 2 -subspace embedding, and that our resulting approximation converges to the optimal solution.

97 MATHEMATICS AND COMPUTING↗

Benchmarking quantum logic operations relative to thresholds for fault tolerance

Contemporary methods for benchmarking noisy quantum processors typically measure average error rates or process infidelities. However, thresholds for fault-tolerant quantum error correction are given in terms of worst-case error rates—defined via the diamond norm—which can differ from average error rates by orders of magnitude. One method for resolving this discrepancy is to randomize the physical implementation of quantum gates, using techniques like randomized compiling (RC). In this work, we use gate set tomography to perform precision characterization of a set of two-qubit logic gates to study RC on a superconducting quantum processor. We find that, under RC, gate errors are accurately described by a stochastic Pauli noise model without coherent errors, and that spatially correlated coherent errors and non-Markovian errors are strongly suppressed. We further show that the average and worst-case error rates are equal for randomly compiled gates, and measure a maximum worst-case error of 0.0197(3) for our gate set. Our results show that randomized benchmarks are a viable route to both verifying that a quantum processor’s error rates are below a fault-tolerance threshold, and to bounding the failure rates of near-term algorithms, if—and only if—gates are implemented via randomization methods which tailor noise.

97 MATHEMATICS AND COMPUTING↗

Ensemble methods for quantification of potassium oxide in ChemCam Mars and laboratory spectra

In this paper we test new approaches for predicting the amount of element oxides in rock samples from the ChemCam instrument suite onboard the NASA Curiosity rover by focusing on K 2 O. Using the expanded dataset compiled by Gasda et al. (2021) with and without the Earth to Mars (E2M and NoE2M) transformation discussed in Clegg et al. (2017) we trained blended submodels using the “double blending” technique and compared these to ensemble methods (Random Forest, ExtraTrees, and Gradient Boosting Regression). We found that ensemble methods performed similar to blended submodels when looking at RMSE-P on the laboratory spectra and provided significant advantages when looking at spectra coming from Mars. For the full model, blended submodels achieved an RMSE-P of 0.62 and 0.60 (E2M and NoE2M respectively) while Gradient Boosting Regression resulted in a slightly improved RMSE-P of 0.59 and 0.60. More importantly, by employing a local RMSE-P estimation technique where model performance is evaluated based on nearby test samples we found that using ensemble methods can lower the quantification limit for K 2 O from the current value of ≈0.6 wt% to ≈0.08 wt% using Extra Trees and Random Forest. This would allow for a much larger range of K 2 O values to be quantified on Mars with greater certainty given that most targets seen on Mars tend to have <1 wt% K2O. Finally, we used both Mean Decrease in Impurity (MDI) and permutation importance techniques to investigate the wavelengths used by the ensemble methods and found that they correspond to known potassium emission lines. This suggests that ensemble methods can provide an easier to train and improved alternative to blended submodels for predicting potassium compositions from Laser Induced Breakdown Spectroscopy (LIBS) data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Single-shot imaging with randomized structured illumination at a free electron laser

Stroboscopic nanoscale imaging with free electron laser light is revolutionizing our understanding of fast dynamics in heterogeneous systems. The short wavelength of X-ray and extreme ultraviolet radiation makes it possible to achieve nanoscale resolution, while resonance with atomic transitions gives access to electronic and magnetic degrees of freedom. Here, we report on our implementation of a recently developed imaging method, randomized probe imaging, at a free electron laser. The advantage of randomized probe imaging over existing methods is its compatibility with extended and strongly scattering samples. Our implementation delivers robust single-shot reconstructions at up to a full-pitch resolution of 400 nm over a field of view with a 40 µm diameter. We also demonstrate single-shot imaging of magnetic domain structures using circular dichroism at resonance, paving the way to future time-resolved studies of magnetic dynamics, shock physics, and the dynamics of collective electronic phases.

47 OTHER INSTRUMENTATION↗

PyTREES

PyTREES (Python tool for Training/Testing Robust Explainable Ensembles on Spectra) is software that implements a data-driven approach to predicting the amount of specific oxides present in materials samples of laser-induced breakdown spectroscopy (LIBS); such as from the ChemCam instrument suite onboard the NASA Curiosity rover. PyTREES is designed to input LIBS data in the format provided by the ChemCam team [1]. PyTREES then applies appropriate pre-processing to this data [2], and implements several regression methods for predicting oxides from spectra. The regression methods include: ensemble methods (random forest, extra trees, and gradient boosting regression) and blended submodels using the “double blending” technique. PyTREES additionally implements methods for quantifying the importance of features in regression model: (1) mean decrease in impurity (MDI) and (2) permutation importance to investigate the wavelengths used by the regression methods. [1] Gasda et al. (2021). Spectrochim Acta B, 181, 106223. [2] Clegg et al. (2017). Spectrochim Acta B , 129, 64–85.

Oyen, Diane↗

Enhanced Boundary Layer Height Detection Using Ceilometer, Surface Meteorology, and Radiation Products With a Random Forest Ensemble Method

This study develops and evaluates a Random Forest (RF) model for estimating planetary boundary layer height (PBLH) using 9 years of data from the Atmospheric Radiation Measurement Southern Great Plains (ARM SGP) user facility, with potential application in the NOAA Surface Radiation (SURFRAD) Network. The model integrates ceilometer, surface meteorology, and radiation measurements, and is trained using thermodynamic PBLH estimates derived from radiosondes. This approach aims to bridge gaps between aerosol-based and thermodynamic-based PBLH estimates. The RF model outperformed traditional methods during daytime and better captured transition periods, demonstrating improved accuracy and robustness. At ARM SGP, it showed a substantial reduction in both bias and RMSE, with a bias near zero (−4.9 m) compared with traditional Haar Wavelet (HW) (70.9 m) and Vaisala BL-View software (124.1 m), and an RMSE of 303.2 m, lower than both BL-View (566.9 m) and HW (404.6 m). During daytime hours, RF consistently outperformed both alternatives, maintaining lower bias and RMSE across all periods. At a second evaluation site, RF achieved the lowest overall RMSE (323.7 m), similar to HW (326.4 m) and significantly better than BL-View (738.3 m). However, all models showed reduced accuracy under stable nighttime conditions, limiting the reliability of PBLH estimates. Key predictors for the model included the lifting condensation level height (LCLH), aerosol gradients, and month for seasonal variability. The study underscores the potential of integrating machine learning with multiple data sets such as surface energy and thermodynamic data to advance PBLH estimation.

boundary layer height↗

A Practical Comparison of Data-Driven Prognostics Methods for Energy Systems

This study explores data-driven prognostics for nuclear power plant (NPP) condensers, focusing on tube fouling. We utilized the Asherah nuclear power plant simulator (ANS) to compare four methods: Random Forest (RF), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory Neural Network (LSTM). By simulating various fouling scenarios in the ANS, we generated data with different degradation rates under transient operations. The models were trained and tested on these data, with performance evaluated visually and numerically including uncertainty assessment. The LSTM model excelled, exhibiting minimal prediction noise and the most accurate remaining useful life estimates across all degradation levels. Its ability to capture long-term dependencies and produce cleaner outputs makes it a strong candidate, although accurate training data across the entire component lifespan are crucial. The RF model emerged as a robust alternative, providing reliable predictions with high confidence. The FCNN and SVR models, while less effective overall, showed potential under specific conditions. FCNN offers a less complex alternative to LSTM and might benefit from larger datasets. SVR excels in precision when the quality of the training data is high. Furthermore, this study highlights the operational benefits of advanced prognostics in the energy sector and emphasizes the need for further research in NPP condenser health management through real-life experiments.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

ML-Based Power System Stability Assessment Considering Network Topology Changes: WECC 20,000+ Bus System Case Study

Modern power grids are fast-changing and thus require real-time monitoring and online stability assessment. With the rapid development of machine learning (ML) techniques, using data-driven models to provide fast and accurate estimations of power system stability marginal information, such as frequency nadir for frequency stability and critical clearing time (CCT) for transient stability, have become possible. However, despite the numerous research on ML-based methods for frequency nadir and CCT prediction, there is limited work on the impact of different network topology changes. Furthermore, most previous studies only focused on small or synthetic systems, and there is a lack of research on actual large power system models. In this paper, the above issues are addressed by studying the actual U.S. Western Electricity Coordinating Council (WECC) system model with more than 20,000 buses. Massive simulations are conducted in PowerWorld Simulator to study the impact of various topology change scenarios on both frequency stability and transient stability. System operating information is extracted from the success dispatch cases of various network topologies to generate a comprehensive dataset for ML-based models. Two ML methods, random forest (RF) and multilayer perceptron (MLP) neural network, are trained and tested for both frequency nadir prediction and CCT prediction. Test results have proven the models are capable of online stability assessment for large power networks such as the WECC system with sufficient accuracy.

critical clearing time↗

Evaluating the performance of random forest and iterative random forest based methods when applied to gene expression data

Gene-to-gene networks, such as Gene Regulatory Networks (GRN) and Predictive Expression Networks (PEN) capture relationships between genes and are beneficial for use in downstream biological analyses. There exists multiple network inference tools to produce these gene-to-gene networks from matrices of gene expression data. Random Forest-Leave One Out Prediction (RF-LOOP) is a method that has been shown to be efficient at producing these gene-to-gene networks, frequently known as GEne Network Inference with Ensemble of trees (GENIE3). Random Forest can be replaced in this process by iterative Random Forest (iRF), which performs variable selection and boosting. Here we validate that iterative Random Forest-Leave One Out Prediction (iRF-LOOP) produces higher quality networks than GENIE3 (RF-LOOP). We use both synthetic and empirical networks from the Dialogue for Reverse Engineering Assessment and Methods (DREAM) Challenges by Sage Bionetworks, as well as two additional empirical networks created from Arabidopsis thaliana and Populus trichocarpa expression data.

59 BASIC BIOLOGICAL SCIENCES↗