Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Evaluation of Data Catalog Software for Hanford Site Environmental Datasets

Environmental information and data underpin achievement of the U.S. Department of Energy (DOE) Office of Environmental Management (EM) mission at the Hanford Site. The Hanford Environmental Data Management (HEDM) Program is the DOE Richland Operations Office (RL) approach to develop and implement a formal program for managing environmental data and the associated records, materials, and systems at the Hanford Site. The current project, contract, organization, and contractor-specific efforts at managing environmental data sets are insufficient to provide orderly, long-term, site-wide access. A vital element to be created within the HEDM program plan is a catalog of data sources, called the Hanford Environmental Information and Data Index (HEIDI), that will enable long-term access and retrievability for the multiple independent sources of data that might otherwise be difficult to discover. This report compares leading open source and commercial data catalog platforms using criteria to assess the functionality needed to develop the HEIDI catalog of Hanford data sources that connects and exchanges data with established Hanford Local Area Network (HLAN) enterprise information technology systems. Proprietary platforms evaluated included ArcGIS Enterprise Sites, Junar, OpenDataSoft, and Socrata, and non-proprietary platforms included Energy Data eXchange (EDX), Comprehensive Knowledge Archive Network (CKAN), and DKAN (a Drupal-based open data portal based on CKAN). Capabilities supporting data discoverability, retrieval, and archival, as well as metadata standard requirements and integration into the HLAN were rated as either failing to meet requirements (F), meeting requirements (M), or exceeding requirements by delivering additional desired features (E). The lowest rating for any capability area was assigned as the overall rating for the platform. These findings enable DOE-RL and the contractors implementing the HEDM plan to focus on candidate tools likely to meet the requirements for implementing HEIDI. All of the platforms receiving an overall rating of ‘F’ were unable to be deployed on Hanford infrastructure or within dedicated cloud resources. A propriety software-as-a-service (SaaS) model of delivering a data catalog (e.g., found in software such as Junar and OpenDataSoft) favors consistency across customers at the expense of customization and configurable roles that are needed for Hanford work. Hosting data on a shared commercial platform places limits on dataset size (maximum of 240 Mb for OpenDataSoft), a significant limitation for HEIDI implementation. EDX, a government data catalog based on CKAN, received the ‘F’ rating due to an inability to incorporate authentication from HLAN into the system. Among platforms rated ‘M’ or ‘E’, only the Socrata platform had a SaaS delivery model. In contrast to other SaaS platforms, Socrata provided custom roles and gateways that allow local datasets to be incorporated into an online catalog. Socrata also complies with the Federal Risk and Authorization Management Program, a significant benefit for cloud-based management of Hanford data. The other platforms rated ‘M’ or ‘E’, ArcGIS Enterprise Sites, CKAN, and DKAN, provide fully self-hosted options, allowing for greater control and flexibility with the HEIDI catalog. These widely used tools have supportive communities of practice, extensive customization options, and demonstrated deployments that provide evidence that they can meet requirements, often deliver additional desired features, and work well with federal government systems. Completely customized alternatives built on a collection of applications were not evaluated because achieving similar performance to CKAN or DKAN requires substantial resources, especially in the absence of the active communities that have grown to support these tools. ArcGIS Enterprise Sites, Socrata, CKAN, and DKAN were evaluated as strong candidates for successful implementation with HEIDI.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Groundwater Recharge Estimated by Land Surface Models: An Evaluation in the Conterminous U.S.

Estimating diffuse recharge of precipitation is fundamental to assessing groundwater sustainability. Diffuse recharge is also the process through which climate and climate change directly affect groundwater. In this study, we evaluated diffuse recharge over the conterminous U.S. simulated by a suite of land surface models (LSMs) that were forced using a common set of meteorological input data. Simulated annual recharge exhibited spatial patterns that were similar among the LSMs, with the highest values in the eastern U.S. and Pacific Northwest. However, the magnitudes of annual recharge varied significantly among the models and were associated with differences in simulated ET, runoff and snow. Evaluation against two independent datasets did not answer the question of whether the ensemble mean performs the best, due to inconsistency between those datasets. The amplitude and timing of seasonal maximum recharge differed among the models, influenced strongly by model physics governing deep soil moisture drainage rates and, in cold regions, snowmelt. Evaluation using in situ soil moisture observations suggested that true recharge peaks 1-3 months later than simulated recharge, indicating systematic biases in simulating deep soil moisture. However, recharge from lateral flows and through preferential flows cannot be inferred from soil moisture data, and the seasonal cycle of simulated groundwater storage actually compared well with in situ groundwater observations. Long-term trends in recharge were not consistently correlated with either precipitation trends or temperature trends. This study highlights the need to employ dynamic flow models in LSMs, among other improvements, to enable more accurate simulation of recharge.

snowmelt↗

Minutes of the CD-ROM Workshop

The workshop described in this document had two goals: (1) to establish guidelines for the CD-ROM as a tool to distribute datasets; and (2) to evaluate current scientific CD-ROM projects as an archive. Workshop attendees were urged to coordinate with European groups to develop CD-ROM, which is already available at low cost in the U.S., as a distribution medium for astronomical datasets. It was noted that NASA has made the CD Publisher at the National Space Science Data Center (NSSDC) available to the scientific community when the Publisher is not needed for NASA work. NSSDC's goal is to provide the Publisher's user with the hardware and software tools needed to design a user's dataset for distribution. This includes producing a master CD and copies. The prerequisite premastering process is described, as well as guidelines for CD-ROM construction. The production of discs was evaluated. CD-ROM projects, guidelines, and problems of the technology were discussed.

King, Joseph H.↗

The Application of Remotely Sensed Data and Models to Benefit Conservation and Restoration Along the Northern Gulf of Mexico Coast

New data, tools, and capabilities for decision making are significant needs in the northern Gulf of Mexico and other coastal areas. The goal of this project is to support NASA s Earth Science Mission Directorate and its Applied Science Program and the Gulf of Mexico Alliance by producing and providing NASA data and products that will benefit decision making by coastal resource managers and other end users in the Gulf region. Data and research products are being developed to assist coastal resource managers adapt and plan for changing conditions by evaluating how climate changes and urban expansion will impact land cover/land use (LCLU), hydrodynamics, water properties, and shallow water habitats; to identify priority areas for conservation and restoration; and to distribute datasets to end-users and facilitating user interaction with models. The proposed host sites for data products are NOAA s National Coastal Data Development Center Regional Ecosystem Data Management, and Mississippi-Alabama Habitat Database. Tools will be available on the Gulf of Mexico Regional Collaborative website with links to data portals to enable end users to employ models and datasets to develop and evaluate LCLU and climate scenarios of particular interest. These data will benefit the Mobile Bay National Estuary Program in ongoing efforts to protect and restore the Fish River watershed and around Weeks Bay National Estuarine Research Reserve. The usefulness of data products and tools will be demonstrated at an end-user workshop.

Quattrochi, Dale↗

Quantum-assisted associative adversarial network: applying quantum annealing in deep learning

Abstract Generative models have the capacity to model and generate new examples from a dataset and have an increasingly diverse set of applications driven by commercial and academic interest. In this work, we present an algorithm for learning a latent variable generative model via generative adversarial learning where the canonical uniform noise input is replaced by samples from a graphical model. This graphical model is learned by a Boltzmann machine which learns low-dimensional feature representation of data extracted by the discriminator. A quantum processor can be used to sample from the model to train the Boltzmann machine. This novel hybrid quantum-classical algorithm joins a growing family of algorithms that use a quantum processor sampling subroutine in deep learning, and provides a scalable framework to test the advantages of quantum-assisted learning. For the latent space model, fully connected, symmetric bipartite and Chimera graph topologies are compared on a reduced stochastically binarized MNIST dataset, for both classical and quantum sampling methods. The quantum-assisted associative adversarial network successfully learns a generative model of the MNIST dataset for all topologies. Evaluated using the Fréchet inception distance and inception score, the quantum and classical versions of the algorithm are found to have equivalent performance for learning an implicit generative model of the MNIST dataset. Classical sampling is used to demonstrate the algorithm on the LSUN bedrooms dataset, indicating scalability to larger and color datasets. Though the quantum processor used here is a quantum annealer, the algorithm is general enough such that any quantum processor, such as gate model quantum computers, may be substituted as a sampler.

Wilson, Max (ORCID:0000000207983391)↗

Evaluation of daily gridded climate products using in situ FLUXNET data and tree growth modeling

Gridded climate data products have facilitated research in climate and ecology by providing meteorological data continuously across large spatial scales. However, the sensitivity of scientific outcomes to dataset choice remains poorly understood, and evaluation using station-based records can favor datasets built heavily on weather stations. Here, we evaluate seven high-resolution daily gridded datasets covering the contiguous United States using independent meteorology from the FLUXNET2015 dataset, with a focus on the implications of dataset choice for process-based tree growth modeling. We find that gridded products tend to capture temperature accurately while consistently overestimating the magnitude and frequency of precipitation and its extremes. Moreover, datasets vary in how they define a ‘day,’ which significantly affects temporal alignment with FLUXNET2015 observations. Despite differences among the datasets, the interannual variability in tree ring simulations is insensitive to dataset choice, likely because daily-scale biases are averaged out through accumulated growth across several months. However, inaccuracies in temperature and precipitation can significantly bias modeled xylem cell production, with systematically higher annual precipitation in the gridded datasets leading to greater xylem production compared to simulations using in situ data. Our results suggest that model applications, especially those that integrate to time scales longer than one day, are likely insensitive to climate dataset choice, but applications that are sensitive to daily climate variations or to absolute climate values need to carefully consider biases in gridded climate products.

54 ENVIRONMENTAL SCIENCES↗

A Publicly Available, Annotated Dataset for Naturalistic Driving Study and Computer Vision Algorithm Development

Oak Ridge National Laboratory developed and implemented a data collection effort to create a dataset for use in evaluating and testing algorithms for analyzing driver behavior under controlled settings for support of the Federal Highway Administration’s Exploratory Advanced Research Program. This collection is called the ORNL Naturalistic Driving Study Sample (ONDSS). The dataset is designed to emulate aspects of the Second Strategic Highway Research Project (SHRP2), which contained a massive naturalistic driving study (NDS) with over 3000 drivers between 2010 and 2013 using their personal vehicles, with over 4300 person-years of data collected [HANKEY].

42 ENGINEERING↗

Automated identification of local contamination in remote atmospheric composition time series

Abstract. Atmospheric observations in remote locations offer a possibility of exploring trace gas and particle concentrations in pristine environments. However, data from remote areas are often contaminated by pollution from local sources. Detecting this contamination is thus a central and frequently encountered issue. Consequently, many different methods exist today to identify local contamination in atmospheric composition measurement time series, but no single method has been widely accepted. In this study, we present a new method to identify primary pollution in remote atmospheric datasets, e.g., from ship campaigns or stations with a low background signal compared to the contaminated signal. The pollution detection algorithm (PDA) identifies and flags periods of polluted data in five steps. The first and most important step identifies polluted periods based on the derivative (time derivative) of a concentration over time. If this derivative exceeds a given threshold, data are flagged as polluted. Further pollution identification steps are a simple concentration threshold filter, a neighboring points filter (optional), a median, and a sparse data filter (optional). The PDA only relies on the target dataset itself and is independent of ancillary datasets such as meteorological variables. All parameters of each step are adjustable so that the PDA can be “tuned” to be more or less stringent (e.g., flag more or fewer data points as contaminated). The PDA was developed and tested with a particle number concentration dataset collected during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition in the central Arctic. Using strict settings, we identified 62 % of the data as influenced by local contamination. Using a second independent particle number concentration dataset also collected during MOSAiC, we evaluated the performance of the PDA against the same dataset cleaned by visual inspection. The two methods agreed in 94 % of the cases. Additionally, the PDA was successfully applied to a trace gas dataset (CO2), also collected during MOSAiC, and to another particle number concentration dataset, collected at the high-altitude background station Jungfraujoch, Switzerland. Thus, the PDA proves to be a useful and flexible tool to identify periods affected by local contamination in atmospheric composition datasets without the need for ancillary measurements. It is best applied to data representing primary pollution. The user-friendly and open-access code enables reproducible application to a wide suite of different datasets. It is available at https://doi.org/10.5281/zenodo.5761101 (Beck et al., 2021).

54 ENVIRONMENTAL SCIENCES↗

Evaluation of a Regional Crop Model Implementation for Sub-National Yield Assessments in Kenya

CONTEXT: Cropping system models can be used to both assess regional food security and to monitor and predict agricultural drought. Agriculture in Kenya is extremely important to both the economy and food security of the country. OBJECTIVE: This study evaluated a regional implementation of a widely used crop model, the Decision Support System for Agrotechnology Transfer (DSSAT), within a coupled modeling framework, the Regional Hydrologic Extremes Assessment System (RHEAS), over Kenya. The goal of this study was to assess the ability of RHEAS to simulate the annual variability of maize yields at the county level and evaluate the uncertainty inherent in the model and inputs. METHODS: The RHEAS system implements a stochastic ensemble approach to account for field scale variabilities in crop management practices and underlying soil and weather conditions. Satellite-derived datasets were used to evaluate the land surface component of the system and seasonally disaggregated yield for 5 years was used to assess the performance of the cropping system model. RESULTS AND CONCLUSIONS: The median correlation between RHEAS and satellite-derived soil moisture and evapotranspiration estimates were 0.78, and 0.51, respectively, indicating that the model is able to capture the key drivers of the hydrological budget. Overall, RHEAS simulated yearly yield variations with a median correlation of 0.7 with reported yields, with the best performance in the short rains season. However, across both seasons, the RHEAS model was positively biased on the order of ~1.6 MT/ha. The overall median unbiased RMSE was 0.66 MT/ha. The RHEAS system shows skill at simulating extreme departures in anomalies, and a majority of the time (62.5%) the reported yields fall within the interquartile range of the simulations. SIGNIFICANCE: One of the most important areas of improvement for the next generation of agricultural data and models is to better understand and communicate the inherent uncertainties. This is especially critical in data-limited regions. Here we present a modeling system and its implementation that begins to address these concerns. We demonstrate the ability to simulate broad trends in yields at the county level for sub-annual yields with skills that commensurate previous national/annual level studies.

Crop model↗

Spatially and Temporally Detailed Water and Carbon Footprints of U.S. Electricity Generation and Use

Electricity generation in the United States entails significant water usage and greenhouse gas emissions. However, accurately estimating these impacts is complex due to the intricate nature of the electric grid and the dynamic electricity mix. Existing methods to estimate the environmental consequences of electricity use often generalize across large regions, neglecting spatial and temporal variations in water usage and emissions. Consequently, electric grid dynamics, such as temporal fluctuations in renewable energy resources, are often overlooked in efforts to mitigate environmental impacts. The U.S. Department of Energy (DOE) has initiated the development of resilient energyshed management systems, requiring detailed information on the local electricity mix and its environmental impacts. This study supports DOE's goal by incorporating geographic and temporal variations in the electricity mix of the local electric grid to better understand the environmental impacts of electricity end users. We offer hourly estimates of the U.S. electricity mix, detailing fuel types, water withdrawal intensity, and water consumption intensity for each grid balancing authority through our publicly accessible tool, the Water Integrated Mapping of Power and Carbon Tracker (Water IMPACT). While our primary focus is on evaluating water intensity factors, our dataset and programming scripts for historical and real-time analysis also include evaluations of carbon dioxide (equivalence) intensity within the same modeling framework. This integrated approach offers a comprehensive understanding of the environmental footprint associated with electricity generation and use, enabling informed decision-making to effectively reduce Scope 2 water usage and emissions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A robust approach to Gaussian process implementation

Abstract. Gaussian process (GP) regression is a flexible modeling technique used to predict outputs and to capture uncertainty in the predictions. However, the GP regression process becomes computationally intensive when the training spatial dataset has a large number of observations. To address this challenge, we introduce a scalable GP algorithm, termed MuyGPs, which incorporates nearest-neighbor and leave-one-out cross-validation during training. This approach enables the evaluation of large spatial datasets with state-of-the-art accuracy and speed in certain spatial problems. Despite these advantages, conventional quadratic loss functions used in the MuyGPs optimization, such as root mean squared error (RMSE), are highly influenced by outliers. We explore the behavior of MuyGPs in cases involving outlying observations and, subsequently, develop a robust approach to handle and mitigate their impact. Specifically, we introduce a novel leave-one-out loss function based on the pseudo-Huber function (LOOPH) that effectively accounts for outliers in large spatial datasets within the MuyGPs framework. Our simulation study shows that the LOOPH loss method maintains accuracy despite outlying observations, establishing MuyGPs as a powerful tool for mitigating unusual observation impacts in the large data regime. In the analysis of US ozone data, MuyGPs provides accurate predictions and uncertainty quantification, demonstrating its utility in managing data anomalies. Through these efforts, we advance the understanding of GP regression in spatial contexts.

Mukangango, Juliette↗

A Sample of NASA Langley Unsteady Pressure Experiments for Computational Aerodynamics Code Evaluation

As computational fluid dynamics methods mature, code development is rapidly transitioning from prediction of steady flowfields to unsteady flows. This change in emphasis offers a number of new challenges to the research community, not the least of which is obtaining detailed, accurate unsteady experimental data with which to evaluate new methods. Researchers at NASA Langley Research Center (LaRC) have been actively measuring unsteady pressure distributions for nearly 40 years. Over the last 20 years, these measurements have focused on developing high-quality datasets for use in code evaluation. This paper provides a sample of unsteady pressure measurements obtained by LaRC and available for government, university, and industry researchers to evaluate new and existing unsteady aerodynamic analysis methods. A number of cases are highlighted and discussed with attention focused on the unique character of the individual datasets and their perceived usefulness for code evaluation. Ongoing LaRC research in this area is also presented.

Schuster, David M.↗

Digital twins and deep learning segmentation of defects in monolayer MX 2 phases

Developing methods to understand and control defect formation in nanomaterials offers a promising route for materials discovery. Monolayer MX 2 phases represent a particularly compelling case for defect engineering of nanomaterials due to the large variability in their physical properties as different defects are introduced into their structure. However, effective identification and quantification of defects remain a challenge even as high-throughput scanning transmission electron microscopy methods improve. This study highlights the benefits of employing first principles calculations to produce digital twins for training deep learning segmentation models for defect identification in monolayer MX 2 phases. Around 600 defect structures were obtained using density functional theory calculations, with each monolayer MX 2 structure being subjected to multislice simulations for the purpose of generating the digital twins. Several deep learning segmentation architectures were trained on this dataset, and their performances evaluated under a variety of conditions such as recognizing defects in the presence of unidentified impurities, beam damage, grain boundaries, and with reduced image quality from low electron doses. Further, this digital twin approach allows benchmarking different deep learning architectures on a theory dataset, which enables the study of defect classification under a broad array of finely controlled conditions. It thus opens the door to resolving the underpinning physical reasons for model shortcomings and potentially chart paths forward for automated discovery of materials defect phases in experiments.

36 MATERIALS SCIENCE↗

Predicting critical heat flux with uncertainty quantification and domain generalization using conditional variational autoencoders and deep neural networks

Deep generative models (DGMs) can generate synthetic data samples that closely resemble the original dataset, addressing data scarcity. In this work, we developed a conditional variational autoencoder (CVAE) to augment critical heat flux (CHF) data used for the 2006 Groeneveld lookup table. To compare with traditional methods, a fine-tuned deep neural network (DNN) regression model was evaluated on the same dataset. Both models achieved small mean absolute relative errors, with the CVAE showing more favorable results. Uncertainty quantification (UQ) was performed using repeated CVAE sampling and DNN ensembling. The DNN ensemble improved performance over the baseline, while the CVAE maintained consistent results with less variability and higher confidence. Both models achieved small errors inside and outside the training domain, with slightly larger errors outside. Altogether, the CVAE performed better than the DNN in predicting CHF and exhibited better uncertainty behavior.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Evaluation of adjoint-based observation impacts as a function of forecast length using an Observing System Simulation Experiment

Adjoints of numerical weather prediction models may be employed for Forecast Sensitivity to Observation (FSO) in order to monitor the contribution of ingested observation data on short-term forecast skill. However, the calculation of short-term forecast error is difficult due to the lack of a truly independent dataset for verification. In an Observing System Simulation Experiment framework, the Nature Run is able to provide a true and complete verification dataset and allows accurate evaluation of short term forecast errors. In this work, an OSSE developed at the National Aeronautics and Space Administration Global Modeling and Assimilation Office is used to explore the impact of observational data on forecasts in the 6 to 48 hour range. An adjoint of the Global Earth Observing System model is employed to compare the observation impacts estimated using both self-analysis verification and the true Nature Run verification. Self-analysis verification is found to inflate the estimated forecast error growth during the early forecast period, resulting in overestimations of observation impacts, particularly in the 6-12 hour forecast range. By 48 hours, the self-analysis verification estimates of forecast error and observation impacts more closely match the true values. The fraction of beneficial observations is also overinflated at short forecast times when self-analysis verification is used. The progression of impacts of an individual observation or data type depends on the character of the growth of the initial condition error that each observation affects.

numerical weather prediction↗