Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Lessons from 18 Years of Hyperspectral Infrared Sounder Data

By the end of 2013 NASA and EUMETSAT will have accumulated more than 11 years of AIRS, 6 years of IASI and one year of CrIS data. All three instruments were nominally specified to support the NWC for short term weather forecasting with a five year lifetime, but continue to exceed the accuracy requirement needed for weather forecasting alone. This allows use of their data for a much broader range of applications, including the calibration of broad-band instruments in space and climate research. We illustrate calibration aspects with examples from AIRS, IASI and CrIS using spatially uniform clear conditions, simultaneous nadir overpasses and random nadir samples. The differences between AIRS, IASI and CrIS for the purpose of weather forecasting are small and we expect that the excellent forecast impact demonstrated by the combination of AIRS and IASI will be continued by the combination of CrIS and IASI. Clear data are useful for calibration, but contain no climate signal. The analysis of random nadir samples from AIRS and CrIS identifies larger biases for observation of extreme conditions, represented by 1% and 99%tile data than for non-extreme observations. This is relevant for climate analysis. Resolution of these differences require further work, since they can complicate the continuation of trends established by AIRS with CrIS data, at least for extrema. The unequaled stability of the AIRS data allows us to evaluate trends using random nadir sampled data. We see an increasing frequency in severe storms over land, a decreasing frequency over ocean. The 11 years of AIRS data are too short to tell if these trends are significant from a climate change viewpoint, or if they are parts of multi-decadal oscillations.

CRIS↗

Efficient Sampling of Complex Interdependent and Multiplex Networks

Efficient sampling of interdependent and multiplex infrastructure networks is critical for effectively applying failure and recovery algorithms in real-world settings, as well as to generate property-preserving reduced-order graph-based ensembles that address topological uncertainties. In this paper, we first explore the performance, i.e. the success in preserving graph properties, of graph sampling algorithms for interdependent and multiplex networks with synthetic and real-world graphs. We simulate sampling algorithms under different parameter settings. These settings include probabilistic graph generators, coupling patterns, and various performance metrics. Our results show that while Random Node and Random Walk sampling algorithms perform best for interdependent networks, Random Edge and Forest Fire sampling algorithms perform best for multiplex networks. Second, we propose and implement a novel similarity-based sampling algorithm for multiplex networks that samples only log(N) number of layers of an N-layer multiplex network while yielding computational savings with performance guarantees. Experimental results show that similarity sampling outperforms complete sampling of all layers while decreasing performance costs from a linear scale to a logarithmic one. Our results also indicate that similarity-based sampling outperforms complete sampling and random selection in nearly all scenarios when tested with real-world data.

Subasi, Omer↗

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

The Completed SDSS-IV extended Baryon Oscillation Spectroscopic Survey: Large-scale structure catalogues for cosmological analysis

ABSTRACT We present large-scale structure catalogues from the completed extended Baryon Oscillation Spectroscopic Survey (eBOSS). Derived from Sloan Digital Sky Survey (SDSS) IV Data Release 16 (DR16), these catalogues provide the data samples, corrected for observational systematics, and random positions sampling the survey selection function. Combined, they allow large-scale clustering measurements suitable for testing cosmological models. We describe the methods used to create these catalogues for the eBOSS DR16 Luminous Red Galaxy (LRG) and Quasar samples. The quasar catalogue contains 343 708 redshifts with 0.8 < z < 2.2 over 4808 deg2. We combine 174 816 eBOSS LRG redshifts over 4242 deg2 in the redshift interval 0.6 < z < 1.0 with SDSS-III BOSS LRGs in the same redshift range to produce a combined sample of 377 458 galaxy redshifts distributed over 9493 deg2. Improved algorithms for estimating redshifts allow that 98 per cent of LRG observations result in a successful redshift, with less than one per cent catastrophic failures (Δz > 1000 km s−1). For quasars, these rates are 95 and 2 per cent (with Δz > 3000 km s−1). We apply corrections for trends between the number densities of our samples and the properties of the imaging and spectroscopic data. For example, the quasar catalogue obtains a χ2/DoF = 776/10 for a null test against imaging depth before corrections and a χ2/DoF= 6/8 after. The catalogues, combined with careful consideration of the details of their construction found here-in, allow companion papers to present cosmological results with negligible impact from observational systematic uncertainties.

79 ASTRONOMY AND ASTROPHYSICS↗

Randomized Adiabatic Quantum Linear Solver Algorithm with Optimal Complexity Scaling and Detailed Running Costs

Solving linear systems of equations is a fundamental problem with a wide variety of applications across many fields of science, and there is increasing effort to develop quantum linear solver algorithms. Subaşı et al. [Phys. Rev. Lett. 122, 060504 (2019)] proposed a randomized algorithm inspired by adiabatic quantum computing, based on a sequence of random Hamiltonian simulation steps, with suboptimal scaling in the condition number 𝜅 of the linear system and the target error 𝜖. Here we go beyond these results in several ways. Firstly, using filtering [Lin and Tong, Quantum 4, 361 (2020)] and Poissonization techniques [Cunningham and Roland, ArXiv:2406.03972 (2024)], the algorithm complexity is improved to the optimal scaling 𝑂⁡(𝜅⁢log (1/𝜖))—an exponential improvement in 𝜖, and a shaving of a log 𝜅 scaling factor in 𝜅. Secondly, the algorithm is further modified to achieve constant factor improvements, which are vital as we progress towards hardware implementations on fault-tolerant devices. We introduce a cheaper randomized walk operator method replacing Hamiltonian simulation—which also removes the need for potentially challenging classical precomputations; randomized routines are sampled over optimized random variables; circuit constructions are improved. We obtain a closed formula rigorously upper bounding the expected number of times one needs to apply a block-encoding of the linear system matrix to output a quantum state encoding the solution to the linear system. The upper bound is 837⁢𝜅 at 𝜖 = 10 −10 for Hermitian matrices.

97 MATHEMATICS AND COMPUTING↗

Model-based quantification of image quality

In 1982, Park and Schowengerdt published an end-to-end analysis of a digital imaging system quantifying three principal degradation components: (1) image blur - blurring caused by the acquisition system, (2) aliasing - caused by insufficient sampling, and (3) reconstruction blur - blurring caused by the imperfect interpolative reconstruction. This analysis, which measures degradation as the square of the radiometric error, includes the sample-scene phase as an explicit random parameter and characterizes the image degradation caused by imperfect acquisition and reconstruction together with the effects of undersampling and random sample-scene phases. In a recent paper Mitchell and Netravelli displayed the visual effects of the above mentioned degradations and presented subjective analysis about their relative importance in determining image quality. The primary aim of the research is to use the analysis of Park and Schowengerdt to correlate their mathematical criteria for measuring image degradations with subjective visual criteria. Insight gained from this research can be exploited in the end-to-end design of optical systems, so that system parameters (transfer functions of the acquisition and display systems) can be designed relative to each other, to obtain the best possible results using quantitative measurements.

Hazra, Rajeeb↗

Basin-Size Mapping: Prediction of Metastable Polymorph Synthesizability Across TaC–TaN Alloys

The sizes of the basins of attraction on the potential energy surface are helpful indicators in determining the experimental synthesizability of metastable phases. In principle, these basins can be controlled with changes in thermodynamic conditions such as composition, pressure, and surface energy. Herein, we use random structure sampling to computationally study how alloying smoothly perturbs basin of attraction sizes. The TaC 1-x N x pseudobinary is an ideal test system given the structural and polymorphic contrast of its parent compounds and their technological relevance as epitaxial substrates for Al 1-x Ga x N. While we find limited thermodynamic stability across all computationally observed phases, random structure sampling shows a significant composition region where the rocksalt basin dominates. As such, we predict the potential for the nonequilibrium synthesis of metastable rocksalt TaC 1-x N x alloys as substrates for Al 1-x Ga x N. At higher nitrogen concentrations, other low-energy metastable polymorphs emerge that continue to retain the hexagonal close packing suitable for III-N growth. Confidence in these trends was established through uncertainty quantification of the basin sizes and energy distributions; such analysis utilized the Beta and Dirichlet distributions. In conclusion, we also find (a) polymorph basin sizes can be rationalized in terms of energetic preferences for different coordination environments; and (b) basin sizes universally shrink with increasing nitrogen content, making the system more prone to amorphous growth.

36 MATERIALS SCIENCE↗

Estimating rates and patterns of diversification with incomplete sampling: a case study in the rosids

Premise Recent advances in generating large‐scale phylogenies enable broad‐scale estimation of species diversification. These now common approaches typically are characterized by (1) incomplete species coverage without explicit sampling methodologies and/or (2) sparse backbone representation, and usually rely on presumed phylogenetic placements to account for species without molecular data. We used empirical examples to examine the effects of incomplete sampling on diversification estimation and provide constructive suggestions to ecologists and evolutionary biologists based on those results. Methods We used a supermatrix for rosids and one well‐sampled subclade (Cucurbitaceae) as empirical case studies. We compared results using these large phylogenies with those based on a previously inferred, smaller supermatrix and on a synthetic tree resource with complete taxonomic coverage. Finally, we simulated random and representative taxon sampling and explored the impact of sampling on three commonly used methods, both parametric (RPANDA and BAMM) and semiparametric (DR). Results We found that the impact of sampling on diversification estimates was idiosyncratic and often strong. Compared to full empirical sampling, representative and random sampling schemes either depressed or inflated speciation rates, depending on methods and sampling schemes. No method was entirely robust to poor sampling, but BAMM was least sensitive to moderate levels of missing taxa. Conclusions We suggest caution against uncritical modeling of missing taxa using taxonomic data for poorly sampled trees and in the use of summary backbone trees and other data sets with high representative bias, and we stress the importance of explicit sampling methodologies in macroevolutionary studies.

59 BASIC BIOLOGICAL SCIENCES↗

X-Ray Diffraction Reference Intensity Ratios of Amorphous and Poorly Crystalline Phases: Implications for CheMin on the Mars Science Laboratory

The CheMin instrument on the Mars Science Laboratory (MSL) rover Curiosity is an X-ray diffraction (XRD) and X-ray fluorescence (XRF) instrument capable of providing the mineralogical and chemical compositions of rocks and soils on the surface of Mars. CheMin uses a microfocus X-ray tube with a Co target, transmission geometry, and an energy-discriminating X-ray sensitive CCD to produce simultaneous 2-D XRD patterns and energy-dispersive X-ray histograms from powdered samples. Piezoelectric vibration of the cell is used to randomize the sample to reduce preferred orientation effects. Instrument details are provided in [1, 2, 3]. Analyses of rock and soil samples by the Mars Exploration Rovers (MER) show nanophase ferric oxide (npOx) is a significant component of the Martian global soil [4] and is thought to be one of the major contributing phases that the Curiosity rover will encounter if a soil sample is analyzed in Gale Crater. Because of the nature of this material, npOx will likely contribute to an X-ray amorphous or short-order component of a XRD pattern measured by the CheMin instrument.

Morris, R. V.↗

Adaptive Conformer Sampling for Property Prediction Using the Conductor-like Screening Model for Real Solvents

The valorization of lignocellulose-derived bioproducts requires effective separation from excessive water. Liquid–liquid extraction is a promising low-energy separation technology, but effective extraction requires solvent selection based on the thermodynamic properties of the bioproduct and solvent components. We propose a computational framework for predicting such properties by developing an adaptive conformer selection approach for use with COSMO-RS (conductor-like screening model for real solvents) calculations. In this framework, molecular dynamics simulations are used to generate many molecular structures (conformers) at representative temperatures in varying solvent environments. Conformers are then clustered based on structural metrics in a low-dimensional space and selected using a mixed-integer quadratic programming problem to iteratively insert a sampled conformer. At each iteration, we determine bioproduct properties using COSMO-RS. Here, we demonstrate the capability of the proposed framework on representative bioproducts to show convergence of the adaptive sampling toward experimentally measured properties with fewer calculations than required by random conformer sampling, enabling the improved screening of solvent systems for liquid-phase separation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Certified randomness using a trapped-ion quantum processor

Although quantum computers can perform a wide range of practically important tasks beyond the abilities of classical computers, realizing this potential remains a challenge. An example is to use an untrusted remote device to generate random bits that can be certified to contain a certain amount of entropy. Certified randomness has many applications but is impossible to achieve solely by classical computation. Here we demonstrate the generation of certifiably random bits using the 56-qubit Quantinuum H2-1 trapped-ion quantum computer accessed over the Internet. Our protocol leverages the classical hardness of recent random circuit sampling demonstrations: a client generates quantum ‘challenge’ circuits using a small randomness seed, sends them to an untrusted quantum server to execute and verifies the results of the server. We analyse the security of our protocol against a restricted class of realistic near-term adversaries. Using classical verification with measured combined sustained performance of 1.1 × 10 18 floating-point operations per second across multiple supercomputers, we certify 71,313 bits of entropy under this restricted adversary and additional assumptions. Our results demonstrate a step towards the practical applicability of present-day quantum computers.

computer science↗

Circulating Tumour Cell Numbers Correlate with Platelet Count and Circulating Lymphocyte Subsets in Men with Advanced Prostate Cancer: Data from the ExPeCT Clinical Trial (CTRIAL-IE 15-21)

Interactions between circulating tumour cells (CTCs) and platelets are thought to inhibit natural killer(NK)-cell-induced lysis. We attempted to correlate CTC numbers in men with advanced prostate cancer with platelet counts and circulating lymphocyte numbers. Sixty-one ExPeCT trial participants, divided into overweight/obese and normal weight groups on the basis of a BMI ≥ 25 or <25, were randomized to participate or not in a six-month exercise programme. Blood samples at randomization, and at three and six months, were subjected to ScreenCell filtration, circulating platelet counts were obtained, and flow cytometry was performed on a subset of samples (n = 29). CTC count positively correlated with absolute total lymphocyte count (r 2 = 0.1709, p = 0.0258) and NK-cell count (r 2 = 0.49, p < 0.0001). There was also a positive correlation between platelet count and CTC count (r 2 = 0.094, p = 0.0001). Correlation was also demonstrated within the overweight/obese group (n = 123, p < 0.0001), the non-exercise group (n = 79, p = 0.001) and blood draw samples lacking platelet cloaking (n = 128, p < 0.0001). By flow cytometry, blood samples from the exercise group (n = 15) had a higher proportion of CD3+ T-lymphocytes (p = 0.0003) and lower proportions of B-lymphocytes (p = 0.0264) and NK-cells (p = 0.015) than the non-exercise group (n = 14). These findings suggest that CTCs engage in complex interactions with the coagulation cascade and innate immune system during intravascular transit, and they present an attractive target for directed therapy at a vulnerable stage in metastasis.

60 APPLIED LIFE SCIENCES↗

Multi-omics Characterization of the Host Response to COVID-19

This project is a multi-disciplinary collaboration between investigators at PNNL with expertise in mass spectrometry (MS)-based omics technology development, omics measurement methods development and application, statistics, machine learning and integration of disparate datasets for a systems-level understanding, and expertise in pathogenic coronaviruses, and investigators at the University of Wisconsin-Madison (UW-Madison) with expertise in pathogenic respiratory viruses (e.g. influenza). The goal of this project is to obtain a comprehensive picture of the human host factors critical for the outcome of SARS-CoV-2 infection. We will generate broad untargeted multi-omics profiles using both state-of-the-art and novel instrumentation and approaches to enable the identification of the molecular mechanisms and host response pathways that impact human COVID-19 outcomes. We anticipate these results will lead to the generation of biomarker panels that are predictive of disease outcomes and mechanistic hypotheses that can be further interrogated in future studies and will provide the basis for vaccine or therapeutic development. To do so, we are obtaining and analyzing blood samples from COVID-19 patients with a range of disease outcomes that were treated at the Center Hospital of the National Center for Global Health and Medicine in Tokyo, Japan and other collaborating hospitals in our network. Specifically, this project will fund proteomics and metabolomics analyses of clinical COVID samples, machine learning-based integration of the data, and pathway-based interpretation of the data. This project was funded in June 2020. In the time span of June to September 2020, the project team developed an analytically and statistically robust analysis plan and made various preparations to facilitate sample receipt from our UW-Madison collaborators. This included blocking and randomization of sample prep orders, ordering of reagents and reference materials, and shipping of materials needed for preparation of the samples under BSL3 conditions to our collaborators at UW-Madison. As of FY21, this project has been picked up via a sponsor, the Naval Medical Research Center, which will cover the remainder of the proposed scope of work.

60 APPLIED LIFE SCIENCES↗

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence↗

A systematic analytical framework for multi-source municipal solid waste characterization for energy recovery

Advancing municipal solid waste (MSW) management from disposal-oriented practices toward circular, value-driven systems requires standardized methodologies capable of identifying material composition and resource recoverable potential at the point of generation. Despite extensive research, MSW characterization remains fragmented due to inconsistences in sampling methodologies, waste sorting categories, and temporal coverage across previous studies which limit cross-site comparability, reproducibility, and constrain the reliable evaluation of potential resource recovery pathways. This lack of consistency has hindered the development of a unified framework for MSW characterization and resource assessment. This study introduces a standardized, field-validated protocol for MSW sampling and composition analysis that ensures consistent, traceable data across diverse waste sources. The protocol integrates randomized spatial sampling, systematic material sorting, and controlled subsampling for multi-site and multi-season field campaigns. Validation included MSW collection from residential, grocery, restaurant, and school MSW streams across five U.S. states, including Maryland, Idaho, Virginia, Ohio, and Mississippi, to demonstrate the protocol’s ability to identify source-based composition patterns relevant to resource recovery applications. Grocery and restaurant streams were dominated by food waste and high-moisture organics, while school waste contained higher paper content and residential waste showed greater heterogeneity. Aggregation into energy-relevant fractions highlighted practical recovery pathways via anaerobic digestion or gasification, supporting data-driven planning, policy, and circular economy strategies for sustainable waste management across waste sources.

09 BIOMASS FUELS↗