Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Optical versus radiographic imaging and tomography: introduction to the ROADS feature issue

Optical imaging is an ancient branch of imaging dating back to thousands of years. Radiographic imaging and tomography (RadIT), including the first use of X-rays by Wilhelm Röntgen, and then, $γ$ -rays, energetic charged particles, neutrons, etc. are about 130 years young. The synergies between optical and radiographic imaging can be cast in the framework of these building blocks: Physics, Sources, Detectors, Methods, and Data Science, as described in Appl. Opt. 61, RDS1 (2022). Optical imaging has expanded to include three-dimensional (3D) tomography (including holography), due in to part the invention of optical (including infrared) lasers. RadIT are intrinsically 3D because of the penetrating power of ionizing radiation. Both optical imaging and tomography (OIT) and RadIT are evolving into even higher dimensional regimes, such as time-resolved tomography (4D) and temporarily and spectroscopically resolved tomography (4D + ). Further advances in OIT and RadIT will continue to be driven by desires for higher information yield, higher resolutions, and higher probability models with reduced uncertainties. Synergies in quantum physics, laser-driven sources, low-cost detectors, data-driven methods, automated processing of data, and artificially intelligent data acquisition protocols will be beneficial to both branches of imaging in many applications. These topics, along with an overview of the Radiography, Applied Optics, and Data Science virtual feature issue, are discussed here.

47 OTHER INSTRUMENTATION↗

LSST Undergraduate Internships at Fermilab

The LSST Data Science for Undergraduate summer internship program focuses on data-driven astronomy for undergraduates at Fermilab’s Cosmic Physics Center. The internship activities focus on the development and implementation of data-driven investigatory techniques that will aid in LSST science, as well as prepare the undergraduates for future work in LSST. In particular, the internship activities include a number of opportunities for undergraduates to learn other skills critical for working on LSST --- data science research techniques, software development and engineering, science communication training.

79 ASTRONOMY AND ASTROPHYSICS↗

Toward Improved Regional Hydrological Model Performance Using State-Of-The-Science Data-Informed Soil Parameters

Accurate soil moisture and streamflow data are an aspirational need of many hydrologically relevant fields. Model simulated soil moisture and streamflow hold promise but models require validation prior to application. Calibration methods are commonly used to improve model fidelity but misrepresentation of the true dynamics remains a challenge. In this study, we leverage soil parameter estimates from the Soil Survey Geographic (SSURGO) database and the probability mapping of SSURGO (POLARIS) to improve the representation of hydrologic processes in the Weather Research and Forecasting Hydrological modeling system (WRF-Hydro) over a central California domain. Our results show WRF-Hydro soil moisture exhibits increased correlation coefficients ( r ), reduced biases, and increased Kling-Gupta Efficiencies (KGEs) across seven in situ soil moisture observing stations after updating the model's soil parameters according to POLARIS. Compared to four well-established soil moisture data sets including Soil Moisture Active Passive data and three Phase 2 North American Land Data Assimilation System land surface models, our POLARIS-adjusted WRF-Hydro simulations produce the highest mean KGE (0.69) across the seven stations. More importantly, WRF-Hydro streamflow fidelity also increases, especially in the case where the model domain is set up with SSURGO-informed total soil thickness. The magnitude and timing of peak flow events are better captured, r increases across nine United States Geological Survey stream gages, and the mean KGE across seven of the nine gages increases from 0.12 to 0.66. Our pre-calibration parameter estimate approach, which is transferable to other spatially distributed hydrological models, can substantially improve a model's performance, helping reduce calibration efforts and computational costs.

54 ENVIRONMENTAL SCIENCES↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Data from: “Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats”

This dataset contains supplementary information for a manuscript describing the ESS-DIVE (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) data repository's community data and metadata reporting formats. The purpose of creating the ESS-DIVE reporting formats was to provide guidelines for formatting some of the diverse data types that can be found in the ESS-DIVE repository. The 6 teams of community partners who developed the reporting formats included scientists and engineers from across the Department of Energy National Lab network. Additionally, during the development process, 247 individuals representing 128 institutions provided input on the formats. The primary files in this dataset are 10 data and metadata crosswalk for ESS-DIVE’s reporting formats (all files ending in _crosswalk.csv). The crosswalks compare elements used in each of the reporting formats to other related standards and data resources (e.g., repositories, datasets, data systems). This dataset also contains additional files recommended by ESS-DIVE’s file-level metadata reporting format. Each data file has an associated dictionary (files ending in _dd.csv) which provide a brief description of each standard or data resource consulted in the data reporting format development process. The flmd.csv file describes each file contained within the dataset.

54 ENVIRONMENTAL SCIENCES↗

Dani Sleight Intern Poster

The NRDS Portal is a login-based data storage solution and science data gateway for researchers to centralize and analyze data before publication. Users can upload, edit, review, and approve their own datasets within the site to eventually be published for public use on the main NRDS site. More development was needed to extend NRDS Portal with new Artificial Intelligence features.

99 - GENERAL AND MISCELLANEOUS↗

SIRIUS: Science-Driven Data Management for Multi-Tiered Storage

The data sets being generated by large applications on very large-scale systems are increasing in both size and complexity. At the same time, there are new ways available to store and access these data sets. The goal in this project is to develop software that applications can use to make use of new and existing storage technologies in more sophisticated ways. One challenge in scientific data management is handling ‘hot’ vs ‘cold’ data. Data that is hot is data that is needed (or will be needed soon) in order for the program to continue progressing, while cold data is either output (and so will not be need further during the life of the program) or will not be needed until significantly later in the program’s run. Hot data should be stored in a way that allows fast access. On most systems, economic factors lead to an inverse relationship between storage performance and storage capacity and so fast access storage is limited. This makes it important to correctly place hot and cold data and avoid cold data unnecessarily consuming precious resources. In this reporting period, we addressed this challenge in various ways and at various levels. Data management frameworks offer only limited control to applications in how data is stored. We have added software capabilities for seamlessly moving data between layers of the storage technology using promote and demote functions to existing software frameworks. This gives direct control to applications in deciding what priority data receives. Additionally, we integrated different storage layer management frameworks in order to allow data to be exchanged and moved between storage layers in a consistent way across the application. Further, applications are not always able to directly decide what storage level makes sense for a given piece of data without an understanding of the underlying storage technologies. Data storage frameworks are often positioned to make these sorts of decisions in service of the application. We have added machine-learning based capabilities to data staging frameworks in order to make intelligent decisions about where data should be stored given learning about patterns in previous usage of similar data.

97 MATHEMATICS AND COMPUTING↗

Carbon Storage Technical Viability Approach (CS TVA): An Integrated Approach for Feasibility and Data Resource Assessment

There is currently a poor understanding and lack of workflow to understand the technical viability of carbon storage spatially. To address this gap, the multi-faceted Carbon Storage Technical Viability Approach (CS TVA) is being developed to incorporate CO2 storage resources, environmental and socio-economic justice (EJ/SJ) factors to enable more comprehensive assessments. The CS TVA includes a (1) matrix framework, (2) an integrated and labeled database, (3) a data availability assessment workflow, and (4) spatial data availability assessment results. This approach leverages spatial and data science analytics to communicate data density, uncertainty, and gaps. The workflow can be applied in whole or in part, based on user needs.

Rodriguez, Neyda Cordero↗

The 2024 “Hacking Limnology” Workshop Series and Virtual Summit: Increasing Inclusion, Participation, and Representation in the Aquatic Sciences

The 4th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) Hacking Limnology Workshop and 5th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 15–19 July 2024. During the week, these joint communities engaged in activities at the intersection of big data, open science, modeling, remote sensing, and the aquatic sciences. The weeklong event, with over 100 aquatic science practitioners and enthusiasts, followed a similar structure to previous years, comprising three days of workshops followed by two days of the virtual summit.

54 ENVIRONMENTAL SCIENCES↗

Subtleties in the trainability of quantum machine learning models

A new paradigm for data science has emerged, with quantum data, quantum models, and quantum computational devices. This field, called quantum machine learning (QML), aims to achieve a speedup over traditional machine learning for data analysis. However, its success usually hinges on efficiently training the parameters in quantum neural networks, and the field of QML is still lacking theoretical scaling results for their trainability. Some trainability results have been proven for a closely related field called variational quantum algorithms (VQAs). While both fields involve training a parametrized quantum circuit, there are crucial differences that make the results for one setting not readily applicable to the other. In this work, we bridge the two frameworks and show that gradient scaling results for VQAs can also be applied to study the gradient scaling of QML models. Our results indicate that features deemed detrimental for VQA trainability can also lead to issues such as barren plateaus in QML. Consequently, our work has implications for several QML proposals in the literature. In addition, we provide theoretical and numerical evidence that QML models exhibit further trainability issues not present in VQAs, arising from the use of a training dataset. We refer to these as dataset-induced barren plateaus. These results are most relevant when dealing with classical data, as here the choice of embedding scheme (i.e., the map between classical data and quantum states) can greatly affect the gradient scaling.

97 MATHEMATICS AND COMPUTING↗

A perspective on Bayesian methods applied to materials discovery and design

For more than two decades, there has been increasing interest in developing frameworks for the accelerated discovery and design of novel materials that could enable promising and transformative technologies. The Integrated Computational Materials Engineering (ICME) program called for integrating computational tools to establish linkages along process-structure-property-performance (PSPP) chains. The Materials Genome Initiative called for integrating experiments and computations within data science frameworks as a strategy to accelerate the materials development cycle. While these frameworks and paradigms have been quite influential, traditional ICME or data science-based approaches tend to have some limitations, mainly when querying the materials space is costly and very little information is available. Bayesian methods are more suitable in this context due to their efficiency gains. To this end, the materials discovery problem is framed as a Bayesian Optimization (BO). Different examples in which BO has been applied to solve materials discovery problems are presented. The methods/examples discussed include BO under model uncertainty, multi-information source BO, multi-objective and multi-constraint BO, and batch BO. Bayesian Materials Discovery is a promising area of research that is likely to become more influential as more attention is put on autonomous materials discovery platforms. Therefore, a discussion is provided on the potential development of such methods to increase the ability of existing platforms in materials discovery. Here, the ultimate goal is to pave the way to autonomous materials discovery.

36 MATERIALS SCIENCE↗

PDFDataExtractor: A Tool for Reading Scientific Text and Interpreting Metadata from the Typeset Literature in the Portable Document Format

The layout of portable document format (PDF) files is constant to any screen, and the metadata therein are latent, compared to mark-up languages such as HTML and XML. No semantic tags are usually provided, and a PDF file is not designed to be edited or its data interpreted by software. However, data held in PDF files need to be extracted in order to comply with opensource data requirements that are now government-regulated. In the chemical domain, related chemical and property data also need to be found, and their correlations need to be exploited to enable data science in areas such as data-driven materials discovery. Such relationships may be realized using text-mining software such as the “chemistry-aware” natural-language-processing tool, ChemDataExtractor; however, this tool has limited data-extraction capabilities from PDF files. This study presents the PDFDataExtractor tool, which can act as a plug-in to ChemDataExtractor. It outperforms other PDF-extraction tools for the chemical literature by coupling its functionalities to the chemical-named entityrecognition capabilities of ChemDataExtractor. The intrinsic PDF-reading abilities of ChemDataExtractor are much improved. The system features a template-based architecture. This enables semantic information to be extracted from the PDF files of scientific articles in order to reconstruct the logical structure of articles. While other existing PDF-extracting tools focus on quantity mining, this template-based system is more focused on quality mining on different layouts. PDFDataExtractor outputs information in JSON and plain text, including the metadata of a PDF file, such as paper title, authors, affiliation, email, abstract, keywords, journal, year, document object identifier (DOI), reference, and issue number. With a self-created evaluation article set, PDFDataExtractor achieved promising precision for all key assessed metadata areas of the document text.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

Multi-modal estimation of physical scale properties

Presentation to Grad students and faculty at the University of Arizona to discuss mathematical and data science problems that we are interested in collaboration on as part of SDRD.

97 MATHEMATICS AND COMPUTING↗

COVID-19 Knowledge Graph -- Dataset for SMCDC 2021 Challenge 2

This repository contains the data for the 2021 Smoky Mountains Computational Sciences Data Challenge (SMCDC21) Challenge 2 -- Finding Novel Links in COVID-19 Knowledge Graph. The total size of all files in this repository is 285MB. Challenge website: https://smc-datachallenge.ornl.gov/2021-challenge-2/ More information about the challenge and the data: https://github.com/ORNL/smcdc-2021-covid-kg

60 APPLIED LIFE SCIENCES↗

Flat-histogram extrapolation as a useful tool in the age of big data

In this work, we review recent work by the authors to revisit the concept of extrapolating thermodynamic properties of classical systems using statistical mechanical principles. Specifically, we discuss how the combination of these principles with biased sampling techniques enables the prediction of free energy landscapes and other detailed information, such as structural properties, of the system in question. Remarkably accurate estimates of physical properties across a broad range of conditions have been achieved using this approach, greatly reducing the number of simulations needed to explore a given system's behaviour. While approximate, these extrapolations significantly amplify the amount of reasonably accurate information that can be extracted from simulations enabling a small set of them to feed data-intensive regression algorithms such as neural networks. Thus, this extrapolation methodology represents a useful tool for performing tasks such as high-throughput screening of physical properties, optimising force field parameters, exploring equilibrium phase behaviour, and enabling theory-guided data science for these systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗