Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Deep Learning without Global Optimization by Random Fourier Neural Networks

Here we introduce a new training algorithm for deep neural networks that utilize random complex exponential activation functions. Our approach employs a Markov chain Monte Carlo sampling procedure to iteratively train network layers, avoiding global and gradient-based optimization while maintaining error control. It consistently attains the theoretical approximation rate for residual networks with complex exponential activation functions, determined by network complexity. Additionally, it enables efficient learning of multiscale and high-frequency features, producing interpretable parameter distributions. Despite using sinusoidal basis functions, we do not observe Gibbs phenomena in approximating discontinuous target functions.

97 MATHEMATICS AND COMPUTING↗

Hierarchical Gaussian Random Field Sampling for Multilevel Markov Chain Monte Carlo: Coupling Stochastic Partial Differential Equation and the Karhunen–Loève Decomposition

This work introduces structure preserving hierarchical decompositions for sampling Gaussian random fields (GRFs) within the context of multilevel Bayesian inference in high-dimensional space. Existing scalable hierarchical sampling methods, such as those based on stochastic partial differential equations (SPDEs), often reduce the dimensionality of the sample space at the cost of accuracy of inference. Other approaches, such that those based on Karhunen-Loève (KL) expansions, offer sample space dimensionality reduction but sacrifice GRF representation accuracy and ergodicity of the Markov chain Monte Carlo (MCMC) sampler and are computationally expensive for high-dimensional problems. The proposed method integrates the dimensionality reduction capabilities of KL expansions with the scalability of SPDE-based sampling, thereby providing a robust, unified framework for high-dimensional uncertainty quantification (UQ) that is scalable and accurate, preserves ergodicity, and offers dimensionality reduction of the sample space. The hierarchy in our multilevel algorithm is derived from the geometric multigrid hierarchy. By constructing a hierarchical decomposition that maintains the covariance structure across the levels in the hierarchy, the approach enables efficient coarse-to-fine sampling while ensuring that all samples are drawn from the desired distribution. The effectiveness of the proposed method is demonstrated on a benchmark subsurface flow problem, demonstrating its effectiveness in improving computational efficiency and statistical accuracy. Furthermore, our proposed technique is more efficient and accurate and displays better convergence properties than existing methods for high-dimensional Bayesian inference problems.

Gaussian random fields↗

Accelerating Random Forest Classification on GPU and FPGA

Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification. In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that for high accuracy targets, our GPU implementation yields 5-9x speedup over CSR, and up to a 2x speedup over cuML.

FPGA, Xilinx FPGA, GPU, Random Forest classificati↗

Solovay-Kitaev Algorithm and Randomized Compilation Data Availability

This zipped folder contains simulation notebooks, simulated data, and experimental data from the QSCOUT trapped-ion device that were used in the publication "Solovay-Kitaev Algorithm and Randomized Compilation" (https://doi.org/10.1103/ll6m-dbl7). The raw data is in the form of measurement outcomes of simple tomographic quantum circuits that were executed on the QSCOUT device and simulated using JAQALPAQ. These data are used to create plots within the jupyter notebooks that were included in the publication.

Quantum benchmarking↗

Comparing Individualized Survival Predictions From Random Survival Forests and Multistate Models in the Presence of Missing Data: A Case Study of Patients With Oropharyngeal Cancer

Background: In recent years, interest in prognostic calculators for predicting patient health outcomes has grown with the popularity of personalized medicine. These calculators, which can inform treatment decisions, employ many different methods, each of which has advantages and disadvantages. Methods: We present a comparison of a multistate model (MSM) and a random survival forest (RSF) through a case study of prognostic predictions for patients with oropharyngeal squamous cell carcinoma. The MSM is highly structured and takes into account some aspects of the clinical context and knowledge about oropharyngeal cancer, while the RSF can be thought of as a black-box non-parametric approach. Key in this comparison are the high rate of missing values within these data and the different approaches used by the MSM and RSF to handle missingness. Results: We compare the accuracy (discrimination and calibration) of survival probabilities predicted by both approaches and use simulation studies to better understand how predictive accuracy is influenced by the approach to (1) handling missing data and (2) modeling structural/disease progression information present in the data. We conclude that both approaches have similar predictive accuracy, with a slight advantage going to the MSM. Conclusions: Although the MSM shows slightly better predictive ability than the RSF, consideration of other differences are key when selecting the best approach for addressing a specific research question. These key differences include the methods’ ability to incorporate domain knowledge, and their ability to handle missing data as well as their interpretability, and ease of implementation. Ultimately, selecting the statistical method that has the most potential to aid in clinical decisions requires thoughtful consideration of the specific goals.

60 APPLIED LIFE SCIENCES↗

A Tutorial for Generating Correlated Random Samples in the Context of Replica Cross Section Data Used in the Propagation of Uncertainty (Second Edition)

The following sample problem write-ups are designed as a tutorial for generating correlated random samples (e.g., replica multi-group cross section data) for use in propagation of uncertainty problems. These problems were set up and solved in MATLAB, but any programming environment with basic statistical functions can be used to generate similar results. Since linear sample problems were chosen, helpful comparisons to the “sandwich” rule propagation of uncertainty are available and were used.

97 MATHEMATICS AND COMPUTING↗

Disulfonated Poly(arylene ether sulfone) Random Copolymers Containing Hierarchical Iptycene Units for Proton Exchange Membranes

Two series of disulfonated iptycene-based poly(arylene ether sulfone) random copolymers, i.e., TRP-BP (triptycene-based) and PENT-BP (pentiptycene-based), were synthesized via condensation polymerization from disulfonated monomer and comonomers to prepare proton exchange membranes (PEMs) for potential applications in electrochemical devices such as fuel cell. To investigate the effect of iptycene units on membrane performance, these copolymers were systematically varied in composition (i.e., iptycene content) and the degree of sulfonation (i.e., 30–50%), which were characterized comprehensively in terms of water uptake, swelling ratio, oxidative stability, thermal and mechanical properties, and proton conductivity at various temperatures. Comparing to copolymers without iptycene units, TRP-BP and PENT-BP ionomers showed greatly enhanced thermal and oxidative stabilities due to strong intra- and inter-molecular supramolecular interactions induced by hierarchical iptycene units. In addition, the introduction of iptycene units in general provides PEMs with exceptional dimensional stability of low volume swelling ratio at high water uptakes, which is ascribed to the supramolecularly interlocked structure as well as high fractional free volume of iptycene-based polymers. It is demonstrated that the combination of high proton conductivity and good membrane dimension stability is the result of the synergistic effects of multiple factors including free volume (iptycene content), sulfonation degree, hydrophobicity, and swelling behavior (supramolecular interactions).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evolution of Robustness in Growing Random Networks

Networks are widely used to model the interaction between individual dynamic systems. In many instances, the total number of units and interaction coupling are not fixed in time, and instead constantly evolve. In networks, this means that the number of nodes and edges both change over time. Various properties of coupled dynamic systems, such as their robustness against noise, essentially depend on the structure of the interaction network. Therefore, it is of considerable interest to predict how these properties are affected when the network grows as well as their relationship to the growth mechanism. Here, we focus on the time evolution of a network’s Kirchhoff index. We derive closed-form expressions for its variation in various scenarios, including the addition of both edges and nodes. For the latter case, we investigate the evolution where single nodes with one or two edges connecting to existing nodes are added recursively to a network. In both cases, we derive the relations between the properties of the nodes to which the new node connects along with the global evolution of network robustness. In particular, we show how different scalings of the Kirchhoff index can be obtained as a function of the number of nodes. We illustrate and confirm this theory via numerical simulations of randomly growing networks.

97 MATHEMATICS AND COMPUTING↗

Analysis of Random Forest Modeling Strategies for Multi-Step Wind Speed Forecasting

Although the random forest (RF) model is a powerful machine learning tool that has been utilized in many wind speed/power forecasting studies, there has been no consensus on optimal RF modeling strategies. This study investigates three basic questions which aim to assist in the discernment and quantification of the effects of individual model properties, namely: (1) using a standalone RF model versus using RF as a correction mechanism for the persistence approach, (2) utilizing a recursive versus direct multi-step forecasting strategy, and (3) training data availability on model forecasting accuracy from one to six hours ahead. These questions are investigated utilizing data from the FINO1 offshore platform and Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) C1 site, and testing results are compared to the persistence method. At FINO1, due to the presence of multiple wind farms and high inter-annual variability, RF is more effective as an error-correction mechanism for the persistence approach. The direct forecasting strategy is seen to slightly outperform the recursive strategy, specifically for forecasts three or more steps ahead. Finally, increased data availability (up to ~8 equivalent years of hourly training data) appears to continually improve forecasting accuracy, although changing environmental flow patterns have the potential to negate such improvement. We hope that the findings of this study will assist future researchers and industry professionals to construct accurate, reliable RF models for wind speed forecasting.

54 ENVIRONMENTAL SCIENCES↗

Sparse and Random Sampling Techniques for High-Resolution, Full-Field, BSS-Based Structural Dynamics Identification from Video

Video-based techniques for identification of structural dynamics have the advantage that they are very inexpensive to deploy compared to conventional accelerometer or strain gauge techniques. When structural dynamics from video is accomplished using full-field, high-resolution analysis techniques utilizing algorithms on the pixel time series such as principal components analysis and solutions to blind source separation the added benefit of high-resolution, full-field modal identification is achieved. An important property of video of vibrating structures is that it is particularly sparse. Typically video of vibrating structures has a dimensionality consisting of many thousands or even millions of pixels and hundreds to thousands of frames. However the motion of the vibrating structure can be described using only a few mode shapes and their associated time series. As a result, emerging techniques for sparse and random sampling such as compressive sensing should be applicable to performing modal identification on video. This work presents how full-field, high-resolution, structural dynamics identification frameworks can be coupled with compressive sampling. The techniques described in this work are demonstrated to be able to recover mode shapes from experimental video of vibrating structures when 70% to 90% of the frames from a video captured in the conventional manner are removed.

47 OTHER INSTRUMENTATION↗

Random Forest Regressor-Based Approach for Detecting Fault Location and Duration in Power Systems

Power system failures or outages due to short-circuits or “faults” can result in long service interruptions leading to significant socio-economic consequences. It is critical for electrical utilities to quickly ascertain fault characteristics, including location, type, and duration, to reduce the service time of an outage. Existing fault detection mechanisms (relays and digital fault recorders) are slow to communicate the fault characteristics upstream to the substations and control centers for action to be taken quickly. Fortunately, due to availability of high-resolution phasor measurement units (PMUs), more event-driven solutions can be captured in real time. In this paper, we propose a data-driven approach for determining fault characteristics using samples of fault trajectories. A random forest regressor (RFR)-based model is used to detect real-time fault location and its duration simultaneously. This model is based on combining multiple uncorrelated trees with state-of-the-art boosting and aggregating techniques in order to obtain robust generalizations and greater accuracy without overfitting or underfitting. Four cases were studied to evaluate the performance of RFR: 1. Detecting fault location (case 1), 2. Predicting fault duration (case 2), 3. Handling missing data (case 3), and 4. Identifying fault location and length in a real-time streaming environment (case 4). A comparative analysis was conducted between the RFR algorithm and state-of-the-art models, including deep neural network, Hoeffding tree, neural network, support vector machine, decision tree, naive Bayesian, and K-nearest neighborhood. Experiments revealed that RFR consistently outperformed the other models in detection accuracy, prediction error, and processing time.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Weather and Random Forest-based Load Profiling Approximation Models and its Transferability across Climate Zones

This study is to provide predictive understanding of the associations of various weather attributes with residential and commercial load profiles, for a variety of climate zones and seasons. In this work, machine learning (ML) approaches were used to identify and quantify the impacts of various weather attributes on residential and commercial electricity demand and its components across the western United States. Performance and transferability of the developed ML models were then evaluated across different temperate zones (e.g., southern, middle, and northern US) and across coastal, mid-continent, and wet zones, with inputs of weather condition data from the National Oceanic and Atmospheric Administration (NOAA) at representative weather stations. The predictive models were developed based on the ranked/screened factors using the regression tree (RT) and random forest (RF) approaches, for five different scenarios (seasons).

load composite, Random Forest, regression tree, lo↗

The Evolution of Randomized Clinical Trial Designs to Assess Therapeutics in Alzheimer Disease

Importance The success of recent randomized clinical trials (RCTs) for Alzheimer disease (AD), particularly those focusing on anti-amyloid therapies, has been discussed at length. However, the evolution of RCT design features for AD that preceded this success remain underexplored. Objective To describe temporal changes in the features of RCT design for interventions in AD. Evidence Review PubMed, Scopus, and Web of Science databases were searched in January 2025 for phase 2 and 3 AD RCTs published between January 1992 and December 2024. RCTs that investigated an intervention for AD, with a placebo or standard-of-care control group, were included. Four assessors independently reviewed full-text articles to capture study characteristics. Main Outcomes and Measures The number of participants and the duration of RCTs as well as the target population, outcomes, and funding were extracted from published reports. These features were analyzed with respect to time using linear regression and χ 2 analyses. Results The study included 203 RCTs with 79 589 participants testing interventions in AD. From 1992 to 2024, the mean sample size increased by 464% for phase 2 RCTs (from 42 to 237), and 50% for phase 3 RCTs (from 632 to 951), while the mean trial duration increased by 188% (from 16 to 46 weeks) for phase 2, and 256% (from 20 to 71 weeks) for phase 3 RCTs. This longer duration of RCTs may be partially attributed by a greater share of disease-modifying rather than symptomatic treatments. Similarly, more recent trials required AD biomarker evidence for enrollment (from 1 of 36 [2.7%] before 2006 to 40 of 76 [52.6%] since 2019). A substantial difference in the type of therapeutics researched was observed, with anti-amyloid and anti-tau RCTs being more likely to be funded by the pharmaceutical industry compared with neurotransmitter or other RCTs (anti-amyloid or anti-tau, 68 of 71 [95.8%]; neurotransmitter, 52 of 69 [77.6%]; other, 33 of 52 [63.5%]). RCT transparency improved, with more frequent data accessibility statements, registered reports, and better reporting on race and ethnicity. Conclusions and Relevance This methodology research of AD RCTs highlights substantial changes in key features of AD clinical trials from 1992 to 2024. AD RCTs have become larger and longer, such that they are powered to detect smaller clinical differences. The increased sample sizes and duration should enable the detection of smaller and more slowly occurring outcomes, which may lead to successful RCTs of therapies with slower and more subtle efficacy.

General & Internal Medicine↗

Nonvolatile Electrochemical Random‐Access Memory under Short Circuit

Abstract Electrochemical random‐access memory (ECRAM) is a recently developed and highly promising analog resistive memory element for in‐memory computing. One longstanding challenge of ECRAM is attaining retention time beyond a few hours. This short retention has precluded ECRAM from being considered for inference classification in deep neural networks, which is likely the largest opportunity for in‐memory computing. In this work, an ECRAM cell with orders of magnitude longer retention than previously achieved is developed, and which is anticipated to exceed ten years at 85 °C. This study hypothesizes that the origin of this exceptional retention is phase separation, which enables the formation of multiple effectively equilibrium resistance states. This work highlights the promises and opportunities to use phase separation to yield ECRAM cells with exceptionally long, and potentially permanent, retention times.

Kim, Diana S.↗

A modified electrolyte non-random two-liquid model with analytical expression for excess enthalpy: Application to the MEA-H 2 O-CO 2 system

We report accurate thermodynamic properties of electrolyte systems are critical for the design and operation of many chemical processes. A comprehensive description of the thermodynamic framework for multi-electrolyte mixed solvent systems is presented, where the parameter structure of the symmetric electrolyte-Non-Random Two Liquid (e-NRTL) model is reformulated and a thermodynamically consistent and analytically derived formulation for the excess enthalpy is developed from the e-NRTL model. The refined parameter structure of the e-NRTL model avoids numerical singularities of the analytical formulation for the excess enthalpy in the absence of ionic species and extends the derived excess enthalpy formulation to non-electrolyte systems. The thermodynamic framework is demonstrated for the MEA-H 2 O-CO 2 case study using experimental data on thermodynamic quantities for the binary MEA-H 2 O system and the ternary MEA-H 2 O-CO 2 system. The model is implemented in Pyomo and will be available for release in the Institute for the Design of Advanced Energy Systems (IDAES) computational platform.

Monoethanolamine↗

Morphological Studies of Solution-Crystallized Thermoplastic Elastomers with Polyethylene Endblocks and a Random-Copolymer Midblock

Styrenic thermoplastic elastomers (TPEs) in the form of triblock copolymers possessing glassy endblocks and a rubbery midblock account for the largest global market of TPEs worldwide, and typically rely on microphase separation of the endblocks and the subsequent formation of rigid microdomains to ensure satisfactory network stabilization. In this work, the morphological characteristics of a relatively new family of crystallizable TPEs that instead consist of polyethylene endblocks and a random-copolymer midblock composed of styrene and (ethylene-co-butylene) moieties are investigated. Copolymer solutions prepared at logarithmic concentrations in a slightly endblock-selective solvent are subjected to crystallization under different time and temperature conditions to ascertain if copolymer self-assembly is directed by endblock crystallization or vice versa. According to transmission electron microscopy, semicrystalline aggregates develop at the lowest solution concentration examined (0.01 wt%), and the size and population of crystals, which dominate the copolymer morphologies, are observed to increase with increasing aging time. Real-space results are correlated with small- and wide-angle X-ray scattering to elucidate the concurrent roles of endblock crystallization and self-assembly of these unique TPEs in solution.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bayesian Calibration of Stochastic Agent Based Model via Random Forest

Agent-based models (ABM) provide an excellent framework for modeling outbreaks and interventions in epidemiology by explicitly accounting for diverse individual interactions and environments. However, these models are usually stochastic and highly parametrized, requiring precise calibration for predictive performance. When considering realistic numbers of agents and properly accounting for stochasticity, this high-dimensional calibration can be computationally prohibitive. This paper presents a random forest-based surrogate modeling technique to accelerate the evaluation of ABMs and demonstrates its use to calibrate an epidemiological ABM named CityCOVID via Markov chain Monte Carlo (MCMC). The technique is first outlined in the context of CityCOVID's quantities of interest, namely hospitalizations and deaths, by exploring dimensionality reduction via temporal decomposition with principal component analysis (PCA) and via sensitivity analysis. The calibration problem is then presented, and samples are generated to best match COVID-19 hospitalization and death numbers in Chicago from March to June in 2020. Further, these results are compared with previous approximate Bayesian calibration (IMABC) results, and their predictive performance is analyzed, showing improved performance with a reduction in computation.

60 APPLIED LIFE SCIENCES↗