Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Multi-omics Characterization of the Host Response to COVID-19

This project is a multi-disciplinary collaboration between investigators at PNNL with expertise in mass spectrometry (MS)-based omics technology development, omics measurement methods development and application, statistics, machine learning and integration of disparate datasets for a systems-level understanding, and expertise in pathogenic coronaviruses, and investigators at the University of Wisconsin-Madison (UW-Madison) with expertise in pathogenic respiratory viruses (e.g. influenza). The goal of this project is to obtain a comprehensive picture of the human host factors critical for the outcome of SARS-CoV-2 infection. We will generate broad untargeted multi-omics profiles using both state-of-the-art and novel instrumentation and approaches to enable the identification of the molecular mechanisms and host response pathways that impact human COVID-19 outcomes. We anticipate these results will lead to the generation of biomarker panels that are predictive of disease outcomes and mechanistic hypotheses that can be further interrogated in future studies and will provide the basis for vaccine or therapeutic development. To do so, we are obtaining and analyzing blood samples from COVID-19 patients with a range of disease outcomes that were treated at the Center Hospital of the National Center for Global Health and Medicine in Tokyo, Japan and other collaborating hospitals in our network. Specifically, this project will fund proteomics and metabolomics analyses of clinical COVID samples, machine learning-based integration of the data, and pathway-based interpretation of the data. This project was funded in June 2020. In the time span of June to September 2020, the project team developed an analytically and statistically robust analysis plan and made various preparations to facilitate sample receipt from our UW-Madison collaborators. This included blocking and randomization of sample prep orders, ordering of reagents and reference materials, and shipping of materials needed for preparation of the samples under BSL3 conditions to our collaborators at UW-Madison. As of FY21, this project has been picked up via a sponsor, the Naval Medical Research Center, which will cover the remainder of the proposed scope of work.

60 APPLIED LIFE SCIENCES↗

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence↗

A systematic analytical framework for multi-source municipal solid waste characterization for energy recovery

Advancing municipal solid waste (MSW) management from disposal-oriented practices toward circular, value-driven systems requires standardized methodologies capable of identifying material composition and resource recoverable potential at the point of generation. Despite extensive research, MSW characterization remains fragmented due to inconsistences in sampling methodologies, waste sorting categories, and temporal coverage across previous studies which limit cross-site comparability, reproducibility, and constrain the reliable evaluation of potential resource recovery pathways. This lack of consistency has hindered the development of a unified framework for MSW characterization and resource assessment. This study introduces a standardized, field-validated protocol for MSW sampling and composition analysis that ensures consistent, traceable data across diverse waste sources. The protocol integrates randomized spatial sampling, systematic material sorting, and controlled subsampling for multi-site and multi-season field campaigns. Validation included MSW collection from residential, grocery, restaurant, and school MSW streams across five U.S. states, including Maryland, Idaho, Virginia, Ohio, and Mississippi, to demonstrate the protocol’s ability to identify source-based composition patterns relevant to resource recovery applications. Grocery and restaurant streams were dominated by food waste and high-moisture organics, while school waste contained higher paper content and residential waste showed greater heterogeneity. Aggregation into energy-relevant fractions highlighted practical recovery pathways via anaerobic digestion or gasification, supporting data-driven planning, policy, and circular economy strategies for sustainable waste management across waste sources.

09 BIOMASS FUELS↗

Masked Symbol Modeling for Demodulation of Oversampled Baseband Communication Signals in Impulsive Noise-Dominated Channels

Recent breakthroughs in natural language processing show that attention mech- anism in Transformer networks, trained via masked-token prediction, enables models to capture the semantic context of the tokens and internalize the grammar of language. While the application of Transformers to communication systems is a burgeoning field, the notion of context within physical waveforms remains under-explored. This paper addresses that gap by re-examining inter-symbol con- tribution (ISC) caused by pulse-shaping overlap. Rather than treating ISC as a nuisance, we view it as a deterministic source of contextual information embedded in oversampled complex baseband signals. We propose Masked Symbol Model- ing (MSM), a framework for the physical (PHY) layer inspired by Bidirectional Encoder Representations from Transformers methodology. In MSM, a subset of symbol-aligned samples is randomly masked, and a Transformer predicts the missing symbol identifiers using the surrounding “in-between” samples. Through this objective, the model learns the latent syntax of complex baseband waveforms. We illustrate MSM’s potential by applying it to the task of demodulating sig- nals corrupted by impulsive noise, where the model infers corrupted segments by leveraging the learned context. Our results suggest a path toward receivers that interpret, rather than merely detect communication signals, opening new avenues for context-aware PHY layer design.

Bedir, Oguz↗

Selection of 3013 Containers for Field Surveillance (FY 2021 Update)

This Update is the ninth in a series of reports that document the binning and sample selection of 3013 containers for the Field Surveillance program as part of the Integrated Surveillance Program. The last Update was in 2016. This Update documents changes to the binning and changes to the random and engineering judgment samples since 2016. This Update also documents field surveillance activities since 2016 and describes plans to complete random sampling in 2025. In addition, this Update describes a new effort to collect surveillance data as part of down blending operations and updates assumptions about Advanced Recovery and Integrated Extraction System (ARIES) 3013 containers.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Critical points of the random cluster model with Newman–Ziff sampling

Here, we present a method for computing transition points of the random cluster model using a generalization of the Newman–Ziff algorithm, a celebrated technique in numerical percolation, to the random cluster model. The new method is straightforward to implement and works for real cluster weight q > 0. Furthermore, results for an arbitrary number of values of q can be found at once within a single simulation. Because the algorithm used to sweep through bond configurations is identical to that of Newman and Ziff, which was conceived for percolation, the method loses accuracy for large lattices when q > 1. However, by sampling the critical polynomial, accurate estimates of critical points in two dimensions can be found using relatively small lattice sizes, which we demonstrate here by computing critical points for non-integer values of q on the square lattice, to compare with the exact solution, and on the unsolved non-planar square matching lattice. The latter results would be much more difficult to obtain using other techniques.

97 MATHEMATICS AND COMPUTING↗

Probabilistic evolution of stochastic dynamical systems: A meso-scale perspective

Stochastic dynamical systems arise naturally across nearly all areas of science and engineering. Typically, a dynamical system model is based on some prior knowledge about the underlying dynamics of interest in which probabilistic features are used to quantify and propagate uncertainties associated with the initial conditions, external excitations, etc. From a probabilistic modeling standing point, two broad classes of methods exist, i.e. macro-scale methods and micro-scale methods. Classically, macro-scale methods such as statistical moments-based strategies are usually too coarse to capture the multi-mode shape or tails of a non-Gaussian distribution. Micro-scale methods such as random samples-based approaches, on the other hand, become computationally very challenging in dealing with high-dimensional stochastic systems. In view of these potential limitations, a meso-scale scheme is proposed here that utilizes a meso-scale statistical structure to describe the dynamical evolution from a probabilistic perspective. The significance of this statistical structure is twofold. First, it can be tailored to any arbitrary random space. Second, it not only maintains the probability evolution around sample trajectories but also requires fewer meso-scale components than the micro-scale samples. To demonstrate the efficacy of the proposed meso-scale scheme, a set of examples of increasing complexity are provided. Connections to the benchmark stochastic models as conservative and Markov models along with practical implementation guidelines are presented.

97 MATHEMATICS AND COMPUTING↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Stochastic AC optimal power flow: A data-driven approach

There is an emerging need for efficient solutions to stochastic AC Optimal Power Flow (AC-OPF) to ensure optimal and reliable grid operations in the presence of increasing demand and generation uncertainty. Herein this paper presents a highly scalable data-driven algorithm for stochastic AC-OPF that has extremely low sample requirement. The novelty behind the algorithm’s performance involves an iterative scenario design approach that merges information regarding constraint violations in the system with data-driven sparse regression. Compared to conventional methods with random scenario sampling, our approach is able to provide feasible operating points for realistic systems with much lower sample requirements. Furthermore, multiple sub-tasks in our approach can be easily paralleled and based on historical data to enhance its performance and application. We demonstrate the computational improvements of our approach through simulations on different test cases in the IEEE PES PGLib-OPF benchmark library.

42 ENGINEERING↗

Microstructure evolution of gadolinium doped cerium oxide under large thermal gradients

In this report the effects of large thermal gradient annealing on the microstructure of 10 mol% gadolinium doped ceria (GDC) were investigated. GDC powder was prepared by solvent deficient method and sintered at 1650 °C for 10 h to achieve dense ceramics with ~8 μm grain size. The densified GDC samples were subsequently annealed using a 60 W infrared laser at over 2100 °C for 1 h under a thermal gradient equivalent to ~0.3–0.5 °C/μm. The post-annealed samples at 2150 °C for 1 h exhibit grains with average length and width of 37 and 28 μm, respectively. Electron backscattered diffraction (EBSD) analysis revealed that the post-annealed sample at 2150 °C consists of grains oriented close to five principal directions (<4 3 10>, <0 0 1> and <13 1 14> on [0 0 1], and <7 6 20> and <7 2 7> on [0 1 0]) within a tolerance angle of ±10°, whereas the grains of the pre-annealed sample are randomly oriented. Gadolinium diffuses 20–30 μm away from the irradiated surface, with the measured composition of regions deeper than 30 μm, Ce 0.86 Gd 0.14 O 1.93 , is close to that of the pre-annealed sample, Ce 0.87 Gd 0.13 O 1.94 . Enhancement of total conductivity of the post-annealed GDC (1.1 × 10 -3 S cm -1 at 500 °C, and 2.1 × 10 -2 S cm -1 at 700 °C) is observed when compared to the pre-annealed GDC (3.1 × 10 -5 S cm -1 at 500 °C, and 1.7 × 10 -3 S cm -1 at 700 °C), and points to the decrease in the grain boundary (GB) resistivity. This could be attributed to both the change in GB area and grain alignment.

36 MATERIALS SCIENCE↗

Avalanches and many-body resonances in many-body localized systems

Here we numerically study both the avalanche instability and many-body resonances in strongly disordered spin chains exhibiting many-body localization (MBL). Finite-size systems behave like MBL within the MBL regimes, which we divide into the asymptotic MBL phase and the finite-size MBL regime; the latter regime is, however, thermal in the limit of large systems and long times. In both Floquet and Hamiltonian models, we identify some landmarks within the MBL regimes. Our first landmark is an estimate of where the MBL phase becomes unstable to avalanches, obtained by measuring the slowest relaxation rate of a finite chain coupled to an infinite bath at one end. Our estimates indicate that the actual MBL-to-thermal phase transition occurs much deeper in the MBL regimes than has been suggested by most previous studies. Our other landmarks involve systemwide many-body resonances: We find that the effective matrix elements producing eigenstates with systemwide many-body resonances are enormously broadly distributed. This broad distribution means that the onset of such resonances in typical samples occurs quite deep in the MBL regimes, and the first such resonances typically involve rare pairs of eigenstates that are farther apart in energy than the minimum gap. Thus we find that the resonance properties define two landmarks that divide the MBL regimes of finite-size systems into three subregimes: (i) at strongest randomness, typical samples do not have any eigenstates that are involved in systemwide many-body resonances; (ii) there is a substantial intermediate subregime where typical samples do have such resonances but the pair of eigenstates with the minimum spectral gap does not, so the size of the minimum gap agrees with expectations from Poisson statistics; and (iii) in the weaker randomness subregime, the minimum gap is larger than predicted by Poisson level statistics because it is involved in a many-body resonance and thus subject to level repulsion. Nevertheless, even in this third subregime, all but a vanishing fraction of eigenstates remain nonresonant and the system thus still appears MBL in most respects. Based on our estimates of the location of the avalanche instability, it might be that the MBL phase is only part of subregime (i) and the other subregimes are entirely in the thermal phase, even though they look localized in most respects, so are in the finite-size MBL regime.

36 MATERIALS SCIENCE↗

SpotSDC: Revealing the Silent Data Corruption Propagation in High-Performance Computing Systems

We report the trend of rapid technology scaling is expected to make the hardware of high-performance computing (HPC) systems more susceptible to computational errors due to random bit flips. Some bit flips may cause a program to crash or have a minimal effect on the output, but others may lead to silent data corruption (SDC), i.e., undetected yet significant output errors. Classical fault injection analysis methods employ uniform sampling of random bit flips during program execution to derive a statistical resiliency profile. However, summarizing such fault injection result with sufficient detail is difficult, and understanding the behavior of the fault-corrupted program is still a challenge. In this article, we introduce SpotSDC, a visualization system to facilitate the analysis of a program's resilience to SDC. SpotSDC provides multiple perspectives at various levels of detail of the impact on the output relative to where in the source code the flipped bit occurs, which bit is flipped, and when during the execution it happens. SpotSDC also enables users to study the code protection and provide new insights to understand the behavior of a fault-injected program. Based on lessons learned, we demonstrate how what we found can improve the fault injection campaign method.

97 MATHEMATICS AND COMPUTING↗

Kinetics of short-range order formation in GeSn alloy: MBE vs CVD

Recently, short-range order (SRO) has attracted significant attention, challenging the conventional view of the atomic positions in alloys being random. Furthermore, the presence of SRO has been predicted to have profound effects on the electronic and topological properties of group-IV alloys, offering a different direction in designing group-IV materials for photoelectronic and quantum devices. However, due to the limited understanding of the formation mechanisms, developing effective methods to manipulate SRO in epitaxy is still challenging. To address this, we propose a mechanism for the GeSn alloy, revealing that surface diffusion plays a key role in SRO formation. Building on this mechanism, we show that the distinct surface conditions in MBE and CVD lead to the formation of SRO with enhanced Sn–Sn pairing in MBE-grown samples, while CVD-grown samples remain random alloys. Furthermore, our findings provide an initial understanding of the kinetic process of SRO formation, providing guidance for the design of experiments to manipulate SRO.

Alloys↗

Classification Analysis of Southwest Pacific Tropical Cyclone Intensity Changes Prior to Landfall

This study evaluates the ability of a random forest classifier to identify tropical cyclone (TC) intensification or weakening prior to landfall over the western region of the Southwest Pacific Ocean (SWPO) basin. For both Australia mainland and SWPO island cases, when a TC first crosses land after spending ≥24 h over the ocean, the closest hour prior to the intersection is considered as the landfall hour. If the maximum wind speed (V max ) at the landfall hour increased or remained the same from the 24-h mark prior to landfall, the TC is labeled as intensifying and if the V max at the landfall hour decreases, the TC is labeled as weakening. Geophysical and aerosol variables closest to the 24 h before landfall hour were collected for each sample. The random forest model with leave-one-out cross validation and the random oversampling example technique was identified as the best-performing classifier for both mainland and island cases. The model identified longitude, initial intensity, and sea skin temperature as the most important variables for the mainland and island landfall classification decisions. Incorrectly classified cases from the test data were analyzed by sorting the cases by their initial intensity hour, landfall hour, monthly distribution, and 24-h intensity changes. TC intensity changes near land strongly impact coastal preparations such as wind damage and flood damage mitigations; hence, this study will contribute to improve identifying and prioritizing prediction of important variables contributing to TC intensity change before landfall.

54 ENVIRONMENTAL SCIENCES↗

OPF-Learn: An Open-Source Framework for Creating Representative AC Optimal Power Flow Datasets: Preprint

Increasing levels of renewable generation motivate a growing interest in data-driven approaches for AC optimal power flow (AC OPF) to manage uncertainty. However, a lack of disciplined dataset creation and benchmarking prohibits useful comparison between approaches in the literature. To instigate confidence, models must be able to reliably predict solutions across a wide range of operating conditions. This paper develops the OPF-Learn package for Julia and Python which uses a computationally efficient approach to create representative datasets that span a wide spectrum of the AC OPF feasible region. Load profiles are uniformly sampled from a convex set that contains the AC OPF feasible set. For each infeasible point found, the convex set is reduced using infeasibility certificates, found by utilizing properties of a relaxed formulation. The framework is shown to generate datasets which are more representative of the entire feasible space versus traditional techniques seen in the literature, improving machine learning model performance.

dataset↗

Non-Boolean quantum amplitude amplification and quantum mean estimation

This paper generalizes the quantum amplitude amplification and amplitude estimation algorithms to work with non-Boolean oracles. The action of a non-Boolean oracle $U_\varphi $ on an eigenstate $\mathinner {|{x}\rangle }$ is to apply a state-dependent phase-shift $\varphi (x)$. Unlike Boolean oracles, the eigenvalues $\exp (i\varphi (x))$ of a non-Boolean oracle are not restricted to be $\pm 1$. Two new oracular algorithms based on such non-Boolean oracles are introduced. The first is the non-Boolean amplitude amplification algorithm, which preferentially amplifies the amplitudes of the eigenstates based on the value of $\varphi (x)$. Starting from a given initial superposition state $\mathinner {|{\psi _0}\rangle }$, the basis states with lower values of $\cos (\varphi )$ are amplified at the expense of the basis states with higher values of $\cos (\varphi )$. The second algorithm is the quantum mean estimation algorithm, which uses quantum phase estimation to estimate the expectation $\mathinner {\langle {\psi _0|U_\varphi |\psi _0}\rangle }$, i.e., the expected value of $\exp (i\varphi (x))$ for a random x sampled by making a measurement on $\mathinner {|{\psi _0}\rangle }$. It is shown that the quantum mean estimation algorithm offers a quadratic speedup over the corresponding classical algorithm. Both algorithms are demonstrated using simulations for a toy example. Potential applications of the algorithms are briefly discussed.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Environment-sensitivity functions for gross primary productivity in light use efficiency models

The sensitivity of photosynthesis to environmental changes is essential for understanding carbon cycle responses to global climate change and for the development of modeling approaches that explains its spatial and temporal variability. We collected a large variety of published sensitivity functions of gross primary productivity (GPP) to different forcing variables to assess the response of GPP to environmental factors. These include the responses of GPP to temperature; vapor pressure deficit, some of which include the response to atmospheric CO 2 concentrations; soil water availability (W); light intensity; and cloudiness. These functions were combined in a full factorial light use efficiency (LUE) model structure, leading to a collection of 5600 distinct LUE models. Each model was optimized against daily GPP and evapotranspiration fluxes from 196 FLUXNET sites and ranked across sites based on a bootstrap approach. The GPP sensitivity to each environmental factor, including CO 2 fertilization, was shown to be significant, and that none of the previously published model structures performed as well as the best model selected. From daily and weekly to monthly scales, the best model's median Nash-Sutcliffe model efficiency across sites was 0.73, 0.79 and 0.82, respectively, but poorer at annual scales (0.23), emphasizing the common limitation of current models in describing the interannual variability of GPP. Although the best global model did not match the local best model at each site, the selection was robust across ecosystem types. The contribution of light saturation and cloudiness to GPP was observed across all biomes (from 23% to 43%). Temperature and W dominates GPP and LUE but responses of GPP to temperature and W are lagged in cold and arid ecosystems, respectively. The findings of this study provide a foundation towards more robust LUE-based estimates of global GPP and may provide a benchmark for other empirical GPP products.

54 ENVIRONMENTAL SCIENCES↗

Spin Dynamics of Quintet and Triplet States Resulting from Singlet Fission in Oriented Terrylenediimide and Quaterrylenediimide Films

Singlet fission in organic semiconductors provides an important opportunity to study high-spin states in electronically coupled chromophores. Photoexcitation of oriented, crystalline films of N,N-bis(pentadec-8-yl)terrylene-3,4:11,12-bis(dicarboximide) (TDI) and N,N-bis (pentadec-8-yl)quaterrylene-3,4:13,14-bis(dicarboximide) (QDI) result in the formation of the correlated triplet pair quintet state, 5 (T 1 T 1 ) and its subsequent dissociation into two triplet (T-1) excitons. Time-resolved electron paramagnetic resonance (TREPR) spectroscopy and X-ray crystallography are used to show that the orientation dependence of 5 (T 1 T 1 ) spin dynamics is retained as it dissociates into two T 1 excitons, while such information is lost in a randomly oriented sample. In addition, the spin dynamics depend on how the molecules are oriented relative to the external applied magnetic field. Dissociation of 5 (T 1 T 1 ) to form two T 1 excitons is more efficient in TDI and leads to a longer T-1 lifetime than in QDI, making TDI a more viable candidate for photovoltaic applications. Since QDI readily forms a highly oriented film, and the lifetime of its 5 (T 1 T 1 ) state is longer than that of TDI, it may be a good candidate for quantum information science applications that require the generation of a quantum- entangled, four-spin state.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗