Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data Inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

LIMS data - Inferred stratospheric distribution of NOx and HOx trace constituents and the calculated odd nitrogen budget

LIMS, SAMS, SBUV and in-situ data have been used to infer species not measured but which are of photochemical interest, e.g., O(3P), O(1D), NO, N2O5, OH, HO2, ClO and HCl. (LIMS = limb infrared monitor of the stratosphere; SAMS = stratospheric and mesospheric sounder; and SBUV = solar backscattered ultraviolet instrument.) Production and loss of odd nitrogen have been calculated and estimates have been made of the odd nitrogen transport due to adiabatically driven circulation derived from LIMS data. Data used from LIMS include O3, NO2, HNO3, H2O and T. CH4 and N2O were taken from SAMS and the UV solar flux from the SBUV instrument. Species were inferred for periods in October, December, March and May. Results for December are discussed. Results indicate: (1) maximum stratospheric odd nitrogen levels of 25 ppbv; (2) evidence of odd nitrogen transport from the mesosphere appearing at 25 km in the wintertime polar latitudes; (3) the polar night build-up of high levels of N2O5 beginning after the autumnal equinox; and (4) the possibility of large downward fluxes of odd nitrogen into the troposphere during the winter at latitudes poleward of 60 degrees.

Callis, L. B.

Applying Gaussian Process Machine Learning and Modern Probabilistic Programming to Satellite Data to Infer CO 2 Emissions

Satellite data provides essential insights into the spatiotemporal distribution of CO 2 concentrations. However, many atmospheric inverse models fail to adequately incorporate the spatial and temporal correlations inherent in satellite observations and often lack rigorous methods for estimating parameters like spatial length scales. We introduce an inference model that processes the spatiotemporal covariance in satellite data and estimates hyperparameters such as covariance length scales. Our approach uses the Gaussian process (GP) machine learning (ML) and modern probabilistic programming languages (PPLs) to perform atmospheric inversions of emissions from satellite data. We develop a GP ML inversion system based on modern PPLs and the GEOS-Chem chemical transport model, simulating atmospheric CO 2 concentrations corresponding to the Orbiting Carbon Observatory-2/3 (OCO-2/3) data for July 2020. In our supervised learning framework, we treat the GEOS-Chem simulated data set as the target, with predictors derived by scaling the target with sector-specific factors hidden from the GP machine. Our results show that the GP model, combined with GPU-enabled PPLs, effectively retrieves true emission scaling factors and infers noise levels concealed within the data. This suggests that our method could be applied over larger areas with more complex covariance structures, enabling comprehensive analysis of the spatiotemporal patterns observed in OCO-2/3 and similar satellite data sets.

54 ENVIRONMENTAL SCIENCES

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING

Data Center High-Temperature Liquid Cooling and Heat Reuse Techno-Economic Study: Preprint

Data centers are energy-intensive facilities with growing demands for efficiency and cost-effective operations. Smaller, more distributed edge inference data centers are expected to proliferate as AI applications require low latency closer to the user of AI tools, which presents a growing opportunity to explore the systems implications of liquid cooling on water and energy use. This study analyzes the implementation of high-temperature liquid cooling systems in a prototypical inference 1-MW data center and explores the potential for heat reuse across varying climates with a goal to optimize energy efficiency, reduce capital and operational costs, and identify opportunities for high-performance cooling and water use reduction infrastructure. This analysis evaluated configurations utilizing a peak day hourly sizing and systems performance spreadsheet to evaluate design and operational conditions from which component sizes, installed cost, operational cost, and performance metrics were determined for the Base case and the Elevated case. The techno-economic analysis included heat reuse applications across a range of heat recovery temperatures and heat rejection options. The analysis shows that high-temperature liquid cooling allows for improved energy efficiency, lower water consumption, and lower capital costs compared to traditional cooling approaches. Transitioning to elevated water inlet/outlet temperatures (50 degrees C/60 degrees C) eliminates the need for chillers, cooling towers, and heat recovery equipment in many scenarios across three distinct climate zones. This results in up to 75% capital cost savings for the cooling and heat recovery equipment, and with significantly reduced water consumption, especially in non-heat reuse applications. Heat generated from data centers can also be repurposed for space heating, domestic hot water, and other applications, and is most cost-effective when data center outlet temperatures exceed 55-60 degrees C.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING

Review of Skin Friction Measurements Including Recent High-Reynolds Number Results from NASA Langley NTF

This paper reviews flat plate skin friction data from early correlations of drag on plates in water to measurements in the cryogenic environment of The NASA Langley National Transonic Facility (NTF) in late 1996. The flat plate (zero pressure gradient with negligible surface curvature) incompressible skin friction at high Reynolds numbers is emphasized in this paper, due to its importance in assessing the accuracy of measurements, and as being important to the aerodynamics of large scale vehicles. A correlation of zero pressure gradient skin friction data minimizing extraneous effects between tests is often used as the first step in the calculation of skin friction in complex flows. Early data compiled by Schoenherr for a range of momentum thickness Reynolds numbers, R(sub Theta) from 860 to 370,000 contained large scatter, but has proved surprisingly accurate in its correlated form. Subsequent measurements in wind tunnels under more carefully controlled conditions have provided inputs to this database, usually to a maximum R(sub Theta) of about 40,000. Data on a large axisymmetric model in the NASA Langley National Transonic Facility extends the upper limit in incompressible R(sub Theta) to 619,800 using the van Driest transformation. Previous data, test techniques, and error sources ar discussed, and the NTF data will be discussed in detail. The NTF Preston tube and Clauser inferred data accuracy is estimated to be within -2 percent of a power-law curve fit, and falls above the Spalding theory by 1 percent at R(sub Theta) of about 600,000.

Watson, Ralph D.

Interhemispheric comparison of atmospheric circulation features as evaluated from Nimbus satellite data

The report includes a complete analyses of O3 data inferred from Nimbus-3 measurements, a discussion of future areas of study, description of the regression and inversion methods developed to infer atmospheric temperature and tropopause characteristics, as well as the plan to process the satellite data for a systematic study of the relative circulation differences between Northern and Southern Hemispheres.

Reiter, E. R.

The use of thermal infrared images in geologic mapping

Thermal infrared image data can be used as an aid to geologic mapping. Broadband thermal data between 8 and 13 microns is used to measure surface temperature, from which surface thermal properties can be inferred. Data from aircraft multispectral scanners at Pisgah, California which include a broadband thermal channel along with several visible and near-IR spectral channels permit better discrimination between rock type units than the same data set without the thermal data. Data from the HCMM satellite and from aircraft thermal scanners also make it possible to monitor moisture changes in Death Valley, California. Multispectral data in the same 8-13 micron wavelength range can be used to discriminate between surface materials with different spectral emission characteristics, as demonstrated with both aircraft scanner and ground spectrometer data.

Kahle, A. B.

Identifying impacts of contact tracing on HIV epidemiological inference from phylogenetic data

Abstract Robust sampling methods are foundational to inferences using phylogenies. Yet the impact of using contact tracing, a type of non-uniform sampling used in public health applications such as infectious disease outbreak investigations, has not been investigated in the molecular epidemiology field. To understand how contact tracing influences a recovered phylogeny, we developed a new simulation tool called SEEPS (Sequence Evolution and Epidemiological Process Simulator) that allows for the simulation of contact tracing and the resulting transmission tree, pathogen phylogeny, and corresponding virus genetic sequences. Importantly, SEEPS takes within-host evolution into account when generating pathogen phylogenies and sequences from transmission histories. Using SEEPS, we demonstrate that contact tracing can significantly impact the structure of the resulting tree, as described by popular tree statistics. Contact tracing generates phylogenies that are less balanced than the underlying transmission process, less representative of the larger epidemiological process, and affects the internal/external branch length ratios that characterize specific epidemiological scenarios. We also examined real data from a 2007–2008 Swedish HIV-1 outbreak and the broader 1998–2010 European HIV-1 epidemic to highlight the differences in contact tracing and expected phylogenies. Aided by SEEPS, we show that the data collection of the Swedish outbreak was strongly influenced by contact tracing even after downsampling, while the broader European Union epidemic showed little evidence of universal contact tracing, agreeing with the known epidemiological information about sampling and spread. Overall, our results highlight the importance of including possible non-uniform sampling schemes when examining phylogenetic trees. For that, SEEPS serves as a useful tool to evaluate such impacts, thereby facilitating better phylogenetic inferences of the characteristics of a disease outbreak. SEEPS is available at https://github.com/MolEvolEpid/SEEPS.

Virology

Crustal magnetization and temperature at depth beneath the Yilgarn block, Western Australia inferred from Magsat data

Variations in crustal magnetization along a seismic section across the Archean Yilgarn block of Western Australia inferred from Magsat data are interpreted as a subtle thermal effect arising from variations in depth to the Curie isotherm. The isotherm lies deep within the mantle of the eastern part of the province, but transects the crust-mantle transition and rises well into the crust on the western side. The model is consistent with heat flow variations along the section line. The mean crustal magnetization implied by the model is approximately 2 A/m. The temperature variation implied by the model is consistent with the hypothesis that the crust-mantle transition seen seismically corresponds to the mafic granulite-eclogite phase transition within a zone of igneous crustal underplating.

Mayhew, M. A.

Identification of the projectile at the Brent crater, and further considerations of projectile types at terrestrial craters

An analysis of impact melt samples from a drill hole at the Brent crater in Ontario for siderophile trace elements indicative of meteoritic contamination, has resulted in 823-857 m-deep basalt melt zone samples enriched in Ir, Os, Pd, Ni, Co, Cr and Se over basement. The abundance pattern suggests a chondritic projectile and, from a Ni-Cr correlation of 10 melt samples, an L or LL chondrite is inferred. Data from the Manicouagan, Mitastin, and Zhamanshin craters are also assessed, and large differences in siderophile element concentrations are found among the tektites which otherwise have similar chemical compositions. There are now four known craters formed by chondrites, with Brent being the smallest among them.

Palme, H.

Maxwell currents under thunderstorms

Time variations observed in thunderstorm electric fields may be interpreted in terms of a total Maxwell current density, varying slowly with time in the intervals between lightning discharges, which can be used to estimate and map thunderstorms. Using the quasi-static behavior of the Maxwell current density, an expression is derived for the field-dependent current density under a thunderstorm during the field recovery following a lightning discharge. Values of air conductivity under the small storm which range from 2 to 6 x 10 to the -13th mho/m are inferred. Data are presented which indicate that the area-average Maxwell current is not usually affected by lightning, and instead varies slowly throughout the evolution of the storm. In light of this, it is suggested that cloud electrification processes probably do not depend on the cloud electric field as much as on the more slowly varying storm dynamics and meteorological structure.

Krider, E. P.

Directional emittance corrections for thermal infrared imaging

A simple measurement technique for measuring the variation of directional emittance of surfaces at various temperatures using commercially available radiometric IR imaging systems was developed and tested. This technique provided the integrated value of directional emittance over the spectral bandwidth of the IR imaging system. The directional emittance of flat black lacquer and red stycast, an epoxy resin, measured using this technique were in good agreement with the predictions of the electromagnetic theory. The data were also in good agreement with directional emittance data inferred from directional reflectance measurements made on a spectrophotometer.

Daryabeigi, Kamran