Engineering PapersSearch

SEARCH · Engineering Papers

Results for “RANDOM SAMPLE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES

Transit Rider/Travel Behavior Inventory Survey - Minneapolis-St. Paul Metro - 1990

The 1990 transit on-board survey aimed to update the 1988 survey, which was conducted as part of the Preliminary Engineering Study for the Hennepin County Light Rail Transit System. The 1988 study received Regional Transit Board funding, and survey results could be applied to a mode split model for projecting ridership on the proposed Light Rail Transit System. However, this survey was designed with the 1990 Travel Behavior Inventory in mind. The intention had been to update the 1988 survey in 1990 to be compatible with the 1990 Travel Behavior Inventory data. The results of the 1990 survey were used primarily to create a table of observed transit trips between each of the 1,200 traffic analysis zones in the region. This trip table was used to calibrate a new mode split model, which estimated future year travel by mode. The 1990 update survey focused on new routes and routes that had changed significantly since 1988. To preserve compatibility with the 1988 survey, the same survey questionnaire card was used, together with the same survey procedures for data collection. The procedures randomly sampled bus patrons during the transit trip, asking key questions about the patron and the transit trip. The survey card was intended for patrons to fill out quickly so it could be completed during the transit trip. The questions focused on conditions that have proven over time to significantly influence ridership. In all, a total of 20,126 valid survey records were processed. Adding surrogate trips, the survey data file is composed of a total of 27,159 trip records . About 10% of the records were filled out by persons who had answered more than one questionnaire.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Testing Classical Properties from Quantum Data

Many properties of Boolean functions can be tested far more efficiently than the function itself can be learned. However, this dramatic advantage often disappears when testers are limited to random samples of ƒ instead of adaptively chosen queries to f. In this work we investigate the quantum version of this restriction: quantum algorithms that test properties of a Boolean function f solely from copies of either the function state |ƒ⟩ ∝ ∑ x |x, ƒ(x)⟩ or the phase state |(-1) ƒ ⟩ ∝ ∑ x (-1) ƒ(x) |x⟩. For monotonicity, symmetry, and triangle-freeness, we show passive quantum testers are unboundedly or super-polynomially better than their classical passive testing counterparts. They are competitive with classic query -based testers in each case. Our new testers use techniques beyond quantum Fourier sampling, and it turns out this is necessary: we show a certain class of bent functions can be tested from 𝒪(1) function states but has a sample complexity lower bound of 2 Ω(n) for any tester relying exclusively on Fourier and classical samples. Our passive quantum testers are competitive with classical query -based testers, but this isn't universal: we exhibit a testing problem that can be solved from 𝒪(1) classical queries but requires Ω(2 n/2 ) function state copies. The Forrelation problem provides a separation of the same magnitude in the opposite direction, so we conclude that quantum data and classical queries are "maximally incomparable" resources for testing. We also begin the study of lower bounds for testing from quantum data. For quantum monotonicity testing, we prove that the ensembles of [Goldreich et al., 2000; Black, 2024], which give exponential lower bounds for classical sample-based testing, do not yield any nontrivial lower bounds for testing from quantum data. New insights specific to quantum data will be required for proving copy complexity lower bounds for testing in this model.

Boolean Functions

A parcel-level evaluation of distributed wind opportunity in the contiguous United States

This study examines the potential for distributed wind (DW) energy across the contiguous United States, leveraging advancements in the National Renewable Energy Laboratory's distributed wind model, dWind. The novel modeling approach described here utilizes a high-resolution dataset and analyzes over 150 million parcels, a significant improvement from prior methods that extrapolated results from a smaller random sample. This achievement is enabled through key model performance improvements, such as transitioning to multiprocessing, which reduces runtime by 97 %. This optimized, high-resolution approach allows the inspection of technology deployment potential and impact on a variety of scales tailored to individual properties and regions. The results here align with prior work showing substantial opportunity for energy generation using DW technologies. Key findings reveal a substantial increase from prior results in estimated technical and economic potential for DW. Metrics tuned to highlight economic potential also show increased incentives supporting rural adoption. Results are spatially aggregated for usability and published via the U.S. Department of Energy Wind Data Portal and a custom scenario visualization platform, aiding policymakers, industry, and property owners in assessing DW viability across various scenarios and spatial scales.

17 WIND ENERGY

Robust Multi-fidelity Bayesian Optimization with Deep Kernel and Partition

Multi-fidelity Bayesian optimization (MFBO) is a powerful approach that utilizes lowfidelity, cost-effective sources to expedite the exploration and exploitation of a high-fidelity objective function. Existing MFBO methods with theoretical foundations either lack justification for performance improvements over single-fidelity optimization or rely on strong assumptions about the relationships between fidelity sources to construct surrogate models and direct queries to low-fidelity sources. To mitigate the dependency on cross-fidelity assumptions while maintaining the advantages of low-fidelity queries, we introduce a random sampling and partition-based MFBO framework with deep kernel learning. This framework is robust to cross-fidelity model misspecification and explicitly illustrates the benefits of low-fidelity queries. Our results demonstrate that the proposed algorithm effectively manages complex cross-fidelity relationships and efficiently optimizes the target fidelity function.

Zhang, Fengxue [University of Chicago, Illinois, U

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Randomized Adiabatic Quantum Linear Solver Algorithm with Optimal Complexity Scaling and Detailed Running Costs

Solving linear systems of equations is a fundamental problem with a wide variety of applications across many fields of science, and there is increasing effort to develop quantum linear solver algorithms. Subaşı et al. [Phys. Rev. Lett. 122, 060504 (2019)] proposed a randomized algorithm inspired by adiabatic quantum computing, based on a sequence of random Hamiltonian simulation steps, with suboptimal scaling in the condition number 𝜅 of the linear system and the target error 𝜖. Here we go beyond these results in several ways. Firstly, using filtering [Lin and Tong, Quantum 4, 361 (2020)] and Poissonization techniques [Cunningham and Roland, ArXiv:2406.03972 (2024)], the algorithm complexity is improved to the optimal scaling 𝑂⁡(𝜅⁢log (1/𝜖))—an exponential improvement in 𝜖, and a shaving of a log 𝜅 scaling factor in 𝜅. Secondly, the algorithm is further modified to achieve constant factor improvements, which are vital as we progress towards hardware implementations on fault-tolerant devices. We introduce a cheaper randomized walk operator method replacing Hamiltonian simulation—which also removes the need for potentially challenging classical precomputations; randomized routines are sampled over optimized random variables; circuit constructions are improved. We obtain a closed formula rigorously upper bounding the expected number of times one needs to apply a block-encoding of the linear system matrix to output a quantum state encoding the solution to the linear system. The upper bound is 837⁢𝜅 at 𝜖 = 10 −10 for Hermitian matrices.

97 MATHEMATICS AND COMPUTING

Basin-Size Mapping: Prediction of Metastable Polymorph Synthesizability Across TaC–TaN Alloys

The sizes of the basins of attraction on the potential energy surface are helpful indicators in determining the experimental synthesizability of metastable phases. In principle, these basins can be controlled with changes in thermodynamic conditions such as composition, pressure, and surface energy. Herein, we use random structure sampling to computationally study how alloying smoothly perturbs basin of attraction sizes. The TaC 1-x N x pseudobinary is an ideal test system given the structural and polymorphic contrast of its parent compounds and their technological relevance as epitaxial substrates for Al 1-x Ga x N. While we find limited thermodynamic stability across all computationally observed phases, random structure sampling shows a significant composition region where the rocksalt basin dominates. As such, we predict the potential for the nonequilibrium synthesis of metastable rocksalt TaC 1-x N x alloys as substrates for Al 1-x Ga x N. At higher nitrogen concentrations, other low-energy metastable polymorphs emerge that continue to retain the hexagonal close packing suitable for III-N growth. Confidence in these trends was established through uncertainty quantification of the basin sizes and energy distributions; such analysis utilized the Beta and Dirichlet distributions. In conclusion, we also find (a) polymorph basin sizes can be rationalized in terms of energetic preferences for different coordination environments; and (b) basin sizes universally shrink with increasing nitrogen content, making the system more prone to amorphous growth.

36 MATERIALS SCIENCE

Certified randomness using a trapped-ion quantum processor

Although quantum computers can perform a wide range of practically important tasks beyond the abilities of classical computers, realizing this potential remains a challenge. An example is to use an untrusted remote device to generate random bits that can be certified to contain a certain amount of entropy. Certified randomness has many applications but is impossible to achieve solely by classical computation. Here we demonstrate the generation of certifiably random bits using the 56-qubit Quantinuum H2-1 trapped-ion quantum computer accessed over the Internet. Our protocol leverages the classical hardness of recent random circuit sampling demonstrations: a client generates quantum ‘challenge’ circuits using a small randomness seed, sends them to an untrusted quantum server to execute and verifies the results of the server. We analyse the security of our protocol against a restricted class of realistic near-term adversaries. Using classical verification with measured combined sustained performance of 1.1 × 10 18 floating-point operations per second across multiple supercomputers, we certify 71,313 bits of entropy under this restricted adversary and additional assumptions. Our results demonstrate a step towards the practical applicability of present-day quantum computers.

computer science

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence

A systematic analytical framework for multi-source municipal solid waste characterization for energy recovery

Advancing municipal solid waste (MSW) management from disposal-oriented practices toward circular, value-driven systems requires standardized methodologies capable of identifying material composition and resource recoverable potential at the point of generation. Despite extensive research, MSW characterization remains fragmented due to inconsistences in sampling methodologies, waste sorting categories, and temporal coverage across previous studies which limit cross-site comparability, reproducibility, and constrain the reliable evaluation of potential resource recovery pathways. This lack of consistency has hindered the development of a unified framework for MSW characterization and resource assessment. This study introduces a standardized, field-validated protocol for MSW sampling and composition analysis that ensures consistent, traceable data across diverse waste sources. The protocol integrates randomized spatial sampling, systematic material sorting, and controlled subsampling for multi-site and multi-season field campaigns. Validation included MSW collection from residential, grocery, restaurant, and school MSW streams across five U.S. states, including Maryland, Idaho, Virginia, Ohio, and Mississippi, to demonstrate the protocol’s ability to identify source-based composition patterns relevant to resource recovery applications. Grocery and restaurant streams were dominated by food waste and high-moisture organics, while school waste contained higher paper content and residential waste showed greater heterogeneity. Aggregation into energy-relevant fractions highlighted practical recovery pathways via anaerobic digestion or gasification, supporting data-driven planning, policy, and circular economy strategies for sustainable waste management across waste sources.

09 BIOMASS FUELS

Masked Symbol Modeling for Demodulation of Oversampled Baseband Communication Signals in Impulsive Noise-Dominated Channels

Recent breakthroughs in natural language processing show that attention mech- anism in Transformer networks, trained via masked-token prediction, enables models to capture the semantic context of the tokens and internalize the grammar of language. While the application of Transformers to communication systems is a burgeoning field, the notion of context within physical waveforms remains under-explored. This paper addresses that gap by re-examining inter-symbol con- tribution (ISC) caused by pulse-shaping overlap. Rather than treating ISC as a nuisance, we view it as a deterministic source of contextual information embedded in oversampled complex baseband signals. We propose Masked Symbol Model- ing (MSM), a framework for the physical (PHY) layer inspired by Bidirectional Encoder Representations from Transformers methodology. In MSM, a subset of symbol-aligned samples is randomly masked, and a Transformer predicts the missing symbol identifiers using the surrounding “in-between” samples. Through this objective, the model learns the latent syntax of complex baseband waveforms. We illustrate MSM’s potential by applying it to the task of demodulating sig- nals corrupted by impulsive noise, where the model infers corrupted segments by leveraging the learned context. Our results suggest a path toward receivers that interpret, rather than merely detect communication signals, opening new avenues for context-aware PHY layer design.

Bedir, Oguz

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING

Kinetics of short-range order formation in GeSn alloy: MBE vs CVD

Recently, short-range order (SRO) has attracted significant attention, challenging the conventional view of the atomic positions in alloys being random. Furthermore, the presence of SRO has been predicted to have profound effects on the electronic and topological properties of group-IV alloys, offering a different direction in designing group-IV materials for photoelectronic and quantum devices. However, due to the limited understanding of the formation mechanisms, developing effective methods to manipulate SRO in epitaxy is still challenging. To address this, we propose a mechanism for the GeSn alloy, revealing that surface diffusion plays a key role in SRO formation. Building on this mechanism, we show that the distinct surface conditions in MBE and CVD lead to the formation of SRO with enhanced Sn–Sn pairing in MBE-grown samples, while CVD-grown samples remain random alloys. Furthermore, our findings provide an initial understanding of the kinetic process of SRO formation, providing guidance for the design of experiments to manipulate SRO.

Alloys

Automated Construction of Artificial Lattice Structures with Designer Electronic States

Manipulating matter with a scanning tunneling microscope (STM) enables the creation of atomically defined artificial structures that host designer quantum states. However, the time-consuming nature of the manipulation process, coupled with the sensitivity of the STM tip, constrains the exploration of diverse configurations and limits the size of the designed features. In this study, we present a reinforcement learning (RL)-based framework for creating artificial structures by spatially manipulating carbon monoxide (CO) molecules on a copper substrate by using the STM tip. The automated workflow combines molecule detection and manipulation, employing deep-learning-based object detection to locate CO molecules and linear assignment algorithms to allocate these molecules to designated target sites. We initially perform molecule maneuvering based on randomized parameter sampling for sample bias, tunneling current set point, and manipulation speed. This data set is then structured into an action trajectory used to train an RL agent. The model is subsequently deployed on the STM for real-time fine-tuning of the manipulation parameters during structure construction. Our approach incorporates path-planning protocols coupled with active drift compensation to enable atomically precise fabrication of structures with significantly reduced human input while realizing larger-scale artificial lattices with the desired electronic properties. Furthermore, using our approach, we demonstrate the automated construction of an extended artificial graphene lattice and confirm the existence of a characteristic Dirac point in its electronic structure. Further challenges regarding the RL-based structural assembly scalability are discussed.

Algorithms