Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

On the Solution of ℓ 0 -Constrained Sparse Inverse Covariance Estimation Problems

The sparse inverse covariance matrix is used to model conditional dependencies between variables in a graphical model to fit a multivariate Gaussian distribution. Estimating the matrix from data are well known to be computationally expensive for large-scale problems. Sparsity is employed to handle noise in the data and to promote interpretability of a learning model. Although the use of a convex ℓ 1 regularizer to encourage sparsity is common practice, the combinatorial ℓ 0 penalty often has more favorable statistical properties. In this paper, we directly constrain sparsity by specifying a maximally allowable number of nonzeros, in other words, by imposing an ℓ 0 constraint. Here, we introduce an efficient approximate Newton algorithm using warm starts for solving the nonconvex ℓ 0 -constrained inverse covariance learning problem. Numerical experiments on standard data sets show that the performance of the proposed algorithm is competitive with state-of-the-art methods.

$\ell_0$-Constrained↗

Probabilistic projections of the Amery Ice Shelf catchment, Antarctica, under conditions of high ice-shelf basal melt

Abstract. Antarctica's Lambert Glacier drains about one-sixth of the ice from the East Antarctic Ice Sheet and is considered stable due to the strong buttressing provided by the Amery Ice Shelf. While previous projections of the sea-level contribution from this sector of the ice sheet have predicted significant mass loss only with near-complete removal of the ice shelf, the ocean warming necessary for this was deemed unlikely. Recent climate projections through 2300 indicate that sufficient ocean warming is a distinct possibility after 2100. This work explores the impact of parametric uncertainty on projections of the response of the Lambert–Amery system (hereafter “the Amery sector”) to abrupt ocean warming through Bayesian calibration of a perturbed-parameter ice-sheet model ensemble. We address the computational cost of uncertainty quantification for ice-sheet model projections via statistical emulation, which employs surrogate models for fast and inexpensive parameter space exploration while retaining critical features of the high-fidelity simulations. To this end, we build Gaussian process (GP) emulators from simulations of the Amery sector at a medium resolution (4–20 km mesh) using the Model for Prediction Across Scales (MPAS)-Albany Land Ice (MALI) model. We consider six input parameters that control basal friction, ice stiffness, calving, and ice-shelf basal melting. From these, we generate 200 perturbed input parameter initializations using space filling Sobol sampling. For our end-to-end probabilistic modeling workflow, we first train emulators on the simulation ensemble and then calibrate the input parameters using observations of the mass balance, grounding line movement, and calving front movement with priors assigned via expert knowledge. Next, we use MALI to project a subset of simulations to 2300 using ocean and atmosphere forcings from a climate model for both low- and high-greenhouse-gas-emission scenarios. From these simulation outputs, we build multivariate emulators by combining GP regression with principal component dimension reduction to emulate multivariate sea-level contribution time series data from the MALI simulations. We then use these emulators to propagate uncertainty from model input parameters to predictions of glacier mass loss through 2300, demonstrating that the calibrated posterior distributions have both greater mass loss and reduced variance compared to the uncalibrated prior distributions. Parametric uncertainty is large enough through about 2130 that the two projections under different emission scenarios are indistinguishable from one another. However, after rapid ocean warming in the first half of the 22nd century, the projections become statistically distinct within decades. Overall, this study demonstrates an efficient Bayesian calibration and uncertainty propagation workflow for ice-sheet model projections and identifies the potential for large sea-level rise contributions from the Amery sector of the Antarctic Ice Sheet after 2100 under high-greenhouse-gas-emission scenarios.

54 ENVIRONMENTAL SCIENCES↗

Uncertain characterization of reservoir fluids due to brittleness of equation of state regression

Equations of state (EoS) play a central role in modeling the phase equilibrium of fluid mixtures. Their parameterization involves fitting a model to experimental data, i.e., solving a nonlinear, non-convex, multivariate optimization problem. The latter requires one to select design variables, domains of definition for each variable, and weights assigned to individual measurements. We demonstrate that subjective choices of an optimization algorithm and an initial guess also impact the regression process. Consequently, EoS predictions are fundamentally uncertain even after the EoS tuning to a limited set of experimental data points. We demonstrate this observation for two hydrocarbon reservoir fluids, in which five properties of the heaviest carbon fraction are treated as design variables. While all the optimization algorithms and initial guesses match experimental data for the gas and liquid properties, the resulting EoS parameterizations lead to dramatically different predictions of the fluid’s thermophysical behavior in the unsampled pressure and temperature regions. In conclusion, we propose the probabilistic treatment of design variables to quantify the predictive uncertainty of the resulting fluid models.

15 GEOTHERMAL ENERGY↗

Capabilities of multivariate Bayesian inference toward seismic hazard assessment

Multivariate Bayesian analysis can bring significant benefits to seismic hazard analysis: Its multivariate feature enables computing scalar and vector hazard without making any approximations; Correlations between intensity measures are implicitly modeled, permitting direct simulation of ground motion selection tools such as the conditional mean spectrum and the generalized conditioning intensity measure; and Its updating feature enables a seamless integration of new ground motion data into the hazard results. Here, we first develop a multivariate Bayesian ground motion model through the NGA-West2 database. The model functional form considers fault-type, magnitude, and distance dependencies, and also the linear and the rock intensity dependent site response. We use a hybrid Markov Chain Monte Carlo sampling to perform Bayesian inference consisting of Gibbs step and a multilevel Metropolis-Hastings step. We then perform several checks on the model and note that its performance is satisfactory. Finally, we illustrate the merits of this multivariate Bayesian analysis, which include: ground motion model updating with ground motion data recorded in the last four years not part of the NGA-West2 database; computation of scalar and vector seismic hazard using the un-updated and updated ground motion models for Los Angeles, CA; and simulation of the conditional mean spectrum under scalar and vector IM conditioning while accounting for different sources of aleatoric and epistemic uncertainties.

58 GEOSCIENCES↗

Field evaluation of semi‐automated moisture estimation from geophysics using machine learning

Geophysical methods can provide three-dimensional (3D), spatially continuous estimates of soil moisture. However, point-to-point comparisons of geophysical properties to measure soil moisture data are frequently unsatisfactory, resulting in geophysics being used for qualitative purposes only. This is because (1) geophysics requires models that relate geophysical signals to soil moisture, (2) geophysical methods have potential uncertainties resulting from smoothing and artifacts introduced from processing and inversion, and (3) results from multiple geophysical methods are not easily combined within a single soil moisture estimation framework. To investigate these potential limitations, an irrigation experiment was performed wherein soil moisture was monitored through time, and several surface geophysical datasets indirectly sensitive to soil moisture were collected before and after irrigation: ground penetrating radar, electrical resistivity tomography (ERT), and frequency domain electromagnetics (FDEM). Data were exported in both raw and processed form, and then snapped to a common 3D grid to facilitate moisture prediction by standard calibration techniques, multivariate regression, and machine learning. A combination of inverted ERT data, raw FDEM, and inverted FDEM data was most informative for predicting soil moisture using a random regression forest model (one-thousand 60/40 training/test cross-validation folds produced root mean squared errors ranging from 0.025–0.046 cm 3 /cm 3 ). This cross-validated model was further supported by a separate evaluation using a test set from a physically separate portion of the study area. Machine learning was conducive to a semi-automated model-selection process that could be used for other sites and datasets to locally improve accuracy.

54 ENVIRONMENTAL SCIENCES↗

MIDAS, prototype Multivariate Interactive Digital Analysis System for large area earth resources surveys. Volume 1: System description

A third-generation, fast, low cost, multispectral recognition system (MIDAS) able to keep pace with the large quantity and high rates of data acquisition from large regions with present and projected sensots is described. The program can process a complete ERTS frame in forty seconds and provide a color map of sixteen constituent categories in a few minutes. A principle objective of the MIDAS program is to provide a system well interfaced with the human operator and thus to obtain large overall reductions in turn-around time and significant gains in throughput. The hardware and software generated in the overall program is described. The system contains a midi-computer to control the various high speed processing elements in the data path, a preprocessor to condition data, and a classifier which implements an all digital prototype multivariate Gaussian maximum likelihood or a Bayesian decision algorithm. Sufficient software was developed to perform signature extraction, control the preprocessor, compute classifier coefficients, control the classifier operation, operate the color display and printer, and diagnose operation.

Christenson, D.↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Spatial estimation from remotely sensed data via empirical Bayes models

Multichannel satellite image data, available as LANDSAT imagery, are recorded as a multivariate time series (four channels, multiple passovers) in two spatial dimensions. The application of parametric empirical Bayes theory to classification of, and estimating the probability of, each crop type at each of a large number of pixels is considered. This theory involves both the probability distribution of imagery data, conditional on crop types, and the prior spatial distribution of crop types. For the latter Markov models indexed by estimable parameters are used. A broad outline of the general theory reveals several questions for further research. Some detailed results are given for the special case of two crop types when only a line transect is analyzed. Finally, the estimation of an underlying continuous process on the lattice is discussed which would be applicable to such quantities as crop yield.

Hill, J. R.↗

EPIsembleVis: A geo-visual analysis and comparison of the prediction ensembles of multiple COVID-19 models

In this work, we present EPIsembleVis, a web-based comparative visual analysis tool for evaluating the consistency of multiple COVID-19 prediction models. Our approach analyzes a collection of COVID-19 predictions from different epidemiological models as an ensemble and utilizes two metrics to quantify model performance. These metrics include (a) prediction uncertainty (represented as the dispersion of predictions in each ensemble) and (b) prediction error (calculated by comparing individual model predictions with the recorded data). Through an interactive visual interface, our approach provides a data-driven workflow for (a) selecting and constructing the COVID-19 model prediction ensemble based on the spatiotemporal overlap of available predictions of multiple epidemiological models, (b) quantifying the model performance using both the uncertainty of each model prediction ensemble, and the error of each ensemble member that represents individual model predictions, and (c) visualizing the spatiotemporal variability in the projection performance of individual models using a suite of novel ensemble visualization techniques, such as the data availability map, a spatiotemporal textured-tile calendar, multivariate rose chart, and time-series leaflet glyph. We demonstrate the capability of our ensemble visual interface through a case study that investigates the performance of weekly COVID-19 predictions, which are provided through the COVID-19 Forecast Hub UMass-Amherst Influenza Forecasting Center of Excellence [47] for the United States and United States Territories. The EPIsembleVis tool is implemented using open-source web technologies and adaptive system design, rendering it interoperable with Elasticsearch and Kibana for automatically ingesting COVID-19 predictions from online repositories, and it is generalizable for analyzing worldwide projections from more epidemiological models.

60 APPLIED LIFE SCIENCES↗

Interdependence in active mobility adoption: Joint modeling and motivational spillover in walking, cycling and bike-sharing

Active mobility offers an array of physical, emotional, and social well-being benefits. However, with the proliferation of the sharing economy, new nonmotorized means of transport are entering the fold, complementing some existing mobility options while competing with others. The purpose of this research study is to investigate the adoption of three active travel modes—namely walking, cycling, and bikesharing—in a joint modeling framework. Here, the analysis is based on an adaptation of the stages of change framework, which originates from the health behavior sciences. Multivariate ordered probit modeling drawing on U.S. survey data provides well-needed insights into individuals’ preparedness to adopt multiple active modes as a function of personal, neighborhood, and psychosocial factors. The research suggests three important findings. (1) The joint model structure confirms interdependence among different active mobility choices. The strongest complementarity is found for walking and cycling adoption. (2) Each mode has a distinctive adoption path with either three or four separate stages. We discuss the implications of derived stage-thresholds and plot adoption contours for selected scenarios. (3) Psychological and neighborhood variables generate more coupling among active modes than individual and household factors. Specifically, identifying strongly with active mobility aspirations, experiences with multimodal travel, possessing better navigational skills, along with supportive local community norms are the factors that appear to drive the joint adoption decisions. This study contributes to the understanding of how decisions within the same functional domain are related and help to design policies that promote active mobility by identifying positive spillovers and joint determinants.

42 ENGINEERING↗

Quantifying Caloric Expenditure During Zero-G Exercise

BACKGROUND: Exercise is a fundamental component of maintaining astronaut health on long-duration space missions, where the microgravity environment poses unique challenges to physiological homeostasis. Accurate quantification of energy expenditure during such exercises is crucial for optimizing nutritional and physical health strategies for spacefarers. OBJECTIVE: This study aims to develop a comprehensive model to estimate caloric expenditure during exercise in a microgravity environment, employing a combination of spirometry, heart rate data, and other relevant parameters. By assessing energy utilization under these conditions, we seek to facilitate enhanced health management protocols for astronauts in space. METHODS & OUTCOMES: A multivariate predictive model will be constructed, utilizing spirometry and heart rate data, coupled with additional physiological and environmental parameters. A systematic approach will be applied to analyze the relationship between these variables and energy expenditure during various exercises. The proposed model will subsequently undergo rigorous validation to ensure accuracy and reliability. This research is expected to yield a precise and reliable predictive model, contributing to improved strategies for exercise prescription and nutritional intake, addressing the unique challenges presented by microgravity environments. We anticipate that our findings will support the development of more effective health maintenance protocols for astronauts during extended space missions, mitigating the adverse effects of space travel on the human body. SIGNIFICANCE: The development of an accurate and adaptable model to quantify caloric expenditure during exercise in space represents a pivotal advancement in space medicine. The insights gained from this study have the potential to inform the design of enhanced health and wellness strategies, ensuring the well-being and operational effectiveness of astronauts in long-duration space missions.

Calorie↗

Appendix B: Principles of computer processing of LANDSAT data

Computer processing facilitates extraction of information from every pixel by executing a variety of functional operations, called processed algorithms, in general or specialized routines. The best results are obtained when data from more than one multispectral band are used together. Multivariate tatistical analysis, computer tape characteristics, processing modes, and a choice of systems (batch or interactive) are discussed. The major operations in computer processing elaborated include: preprocessing, enhancement, effects of rationing, and classification. Techniques for multisource data correlation are considered with emphasis on geobased systems.

Source record↗

Estimating vegetation coverage in St. Joseph Bay, Florida with an airborne multispectral scanner

A four-channel multispectral scanner (MSS) carried aboard an aircraft was used to collect data along several flight paths over St. Joseph Bay, FL. Various classifications of benthic features were defined from the results of ground-truth observations. The classes were statistically correlated with MSS channel signal intensity using multivariate methods. Application of the classification measures to the MSS data set allowed computer construction of a detailed map of benthic features of the bay. Various densities of segrasses, various bottom types, and algal coverage were distinguished from water of various depths. The areal vegetation coverage of St. Joseph Bay was not significantly different from the results of a survey conducted six years previously, suggesting that seagrasses are a very stable feature of the bay bottom.

Savastano, K. J.↗

Nonparametric analysis of Minnesota spruce and aspen tree data and LANDSAT data

The application of nonparametric methods in data-intensive problems faced by NASA is described. The theoretical development of efficient multivariate density estimators and the novel use of color graphics workstations are reviewed. The use of nonparametric density estimates for data representation and for Bayesian classification are described and illustrated. Progress in building a data analysis system in a workstation environment is reviewed and preliminary runs presented.

Scott, D. W.↗

Sensitive Detection of Structural Differences using a Statistical Framework for Comparative Crystallography

Chemical and conformational changes underlie the functional cycles of proteins. Comparative crystallography can reveal these changes over time, over ligands, and over chemical and physical perturbations in atomic detail. A key difficulty, however, is that the resulting observations must be placed on the same scale by correcting for experimental factors. We recently introduced a Bayesian framework for correcting (scaling) X-ray diffraction data by combining deep learning with statistical priors informed by crystallographic theory. To scale comparative crystallography data, we here combine this framework with a multivariate statistical theory of comparative crystallography. By doing so, we find strong improvements in the detection of protein dynamics, element-specific anomalous signal, and the binding of drug fragments.

Hekstra, Doeke R. [Harvard Univ., Cambridge, MA (U↗

WELLS Interactive Application

The Wellbore Exploration and Location Logistic System (WELLS) Interactive Application is an interactive tool to enable easy exploration and visualization of the living national wellbore database (WELLS Database (https://edx.netl.doe.gov/dataset/wells_database)). The tool and underlying database were created and are maintained by the National Energy Technology Laboratory (NETL), providing visualization of the more than six million public wellbore records from more than 65 authoritative state, federal, and tribal resources. The WELLS Interactive Application serves up wellbore data from oil, gas, underground injection, research, geothermal, geotechnical, groundwater, and other types of wells in a single, standardized, unified system. In addition to the surface location of these wells, the underlying database combines select key attributes for features such as well age, depth, and operating status. The system also provides users with references back to the original sources used in this unified platform. The underlying data can be accessed through the WELLS Database: https://edx.netl.doe.gov/dataset/wells_database Additional Information: The WELLS Interactive Application (formerly titled CO2-Locate) enables visualization and access to the public wellbore records through an intuitive web-based mapping tool. The WELLS Interactive Application was designed to help users visualize, query, analyze, and download wellbore records. Public wellbore points are included as a layer in the Map page, called Public Wells. Additionally, a multivariate hexagon grid summarizing well density from proprietary well data, called Well Density, is included to identify data gaps between the public and proprietary well data. Filtering functionalities in the tool allow these two layers to be spatially filtered by state, county, or basin as well as by status, type, true vertical depth, and spud year. The WELLS Interactive Application also contains a Near Me tool can be used to search and explore wellbore data within a user-defined distance of a specified location on the map, which can also be downloaded. The Query tool allows users to query the selected or filtered wells in the Public Wells layer and export the data. For additional information on these tool functionalities, see the help documentation on the About page of the tool. Notes for Consideration: The Well Density layer provided in this application is derived from proprietary wellbore data, the records of which do not always contain values for key features (status, type, true vertical depth, or spud year). Therefore, data might not be available when layers are queried for all filter combinations. Additionally, visualizing layers and applying filters may take additional time to load (i.e., draw on the map) due to the large size of the data.

ccs↗