Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data sciences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The 2024 “Hacking Limnology” Workshop Series and Virtual Summit: Increasing Inclusion, Participation, and Representation in the Aquatic Sciences

The 4th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) Hacking Limnology Workshop and 5th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 15–19 July 2024. During the week, these joint communities engaged in activities at the intersection of big data, open science, modeling, remote sensing, and the aquatic sciences. The weeklong event, with over 100 aquatic science practitioners and enthusiasts, followed a similar structure to previous years, comprising three days of workshops followed by two days of the virtual summit.

54 ENVIRONMENTAL SCIENCES↗

Subtleties in the trainability of quantum machine learning models

A new paradigm for data science has emerged, with quantum data, quantum models, and quantum computational devices. This field, called quantum machine learning (QML), aims to achieve a speedup over traditional machine learning for data analysis. However, its success usually hinges on efficiently training the parameters in quantum neural networks, and the field of QML is still lacking theoretical scaling results for their trainability. Some trainability results have been proven for a closely related field called variational quantum algorithms (VQAs). While both fields involve training a parametrized quantum circuit, there are crucial differences that make the results for one setting not readily applicable to the other. In this work, we bridge the two frameworks and show that gradient scaling results for VQAs can also be applied to study the gradient scaling of QML models. Our results indicate that features deemed detrimental for VQA trainability can also lead to issues such as barren plateaus in QML. Consequently, our work has implications for several QML proposals in the literature. In addition, we provide theoretical and numerical evidence that QML models exhibit further trainability issues not present in VQAs, arising from the use of a training dataset. We refer to these as dataset-induced barren plateaus. These results are most relevant when dealing with classical data, as here the choice of embedding scheme (i.e., the map between classical data and quantum states) can greatly affect the gradient scaling.

97 MATHEMATICS AND COMPUTING↗

A perspective on Bayesian methods applied to materials discovery and design

For more than two decades, there has been increasing interest in developing frameworks for the accelerated discovery and design of novel materials that could enable promising and transformative technologies. The Integrated Computational Materials Engineering (ICME) program called for integrating computational tools to establish linkages along process-structure-property-performance (PSPP) chains. The Materials Genome Initiative called for integrating experiments and computations within data science frameworks as a strategy to accelerate the materials development cycle. While these frameworks and paradigms have been quite influential, traditional ICME or data science-based approaches tend to have some limitations, mainly when querying the materials space is costly and very little information is available. Bayesian methods are more suitable in this context due to their efficiency gains. To this end, the materials discovery problem is framed as a Bayesian Optimization (BO). Different examples in which BO has been applied to solve materials discovery problems are presented. The methods/examples discussed include BO under model uncertainty, multi-information source BO, multi-objective and multi-constraint BO, and batch BO. Bayesian Materials Discovery is a promising area of research that is likely to become more influential as more attention is put on autonomous materials discovery platforms. Therefore, a discussion is provided on the potential development of such methods to increase the ability of existing platforms in materials discovery. Here, the ultimate goal is to pave the way to autonomous materials discovery.

36 MATERIALS SCIENCE↗

PDFDataExtractor: A Tool for Reading Scientific Text and Interpreting Metadata from the Typeset Literature in the Portable Document Format

The layout of portable document format (PDF) files is constant to any screen, and the metadata therein are latent, compared to mark-up languages such as HTML and XML. No semantic tags are usually provided, and a PDF file is not designed to be edited or its data interpreted by software. However, data held in PDF files need to be extracted in order to comply with opensource data requirements that are now government-regulated. In the chemical domain, related chemical and property data also need to be found, and their correlations need to be exploited to enable data science in areas such as data-driven materials discovery. Such relationships may be realized using text-mining software such as the “chemistry-aware” natural-language-processing tool, ChemDataExtractor; however, this tool has limited data-extraction capabilities from PDF files. This study presents the PDFDataExtractor tool, which can act as a plug-in to ChemDataExtractor. It outperforms other PDF-extraction tools for the chemical literature by coupling its functionalities to the chemical-named entityrecognition capabilities of ChemDataExtractor. The intrinsic PDF-reading abilities of ChemDataExtractor are much improved. The system features a template-based architecture. This enables semantic information to be extracted from the PDF files of scientific articles in order to reconstruct the logical structure of articles. While other existing PDF-extracting tools focus on quantity mining, this template-based system is more focused on quality mining on different layouts. PDFDataExtractor outputs information in JSON and plain text, including the metadata of a PDF file, such as paper title, authors, affiliation, email, abstract, keywords, journal, year, document object identifier (DOI), reference, and issue number. With a self-created evaluation article set, PDFDataExtractor achieved promising precision for all key assessed metadata areas of the document text.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

Multi-modal estimation of physical scale properties

Presentation to Grad students and faculty at the University of Arizona to discuss mathematical and data science problems that we are interested in collaboration on as part of SDRD.

97 MATHEMATICS AND COMPUTING↗

COVID-19 Knowledge Graph -- Dataset for SMCDC 2021 Challenge 2

This repository contains the data for the 2021 Smoky Mountains Computational Sciences Data Challenge (SMCDC21) Challenge 2 -- Finding Novel Links in COVID-19 Knowledge Graph. The total size of all files in this repository is 285MB. Challenge website: https://smc-datachallenge.ornl.gov/2021-challenge-2/ More information about the challenge and the data: https://github.com/ORNL/smcdc-2021-covid-kg

60 APPLIED LIFE SCIENCES↗

Comparisons of the v11.1 Orbiting Carbon Observatory‐2 (OCO‐2) X CO2 Measurements With GGG2020 TCCON

The Orbiting Carbon Observatory 2 (OCO-2) is NASA's first Earth observation satellite mission dedicated to studying the sources and sinks of carbon dioxide (CO 2 ) on a global scale. The observations of reflected sunlight are inverted in a retrieval algorithm to produce estimates of the dry air mole-fractions of CO 2 (X CO2 ). The OCO-2 Level 2 data release, version 11.1 (v11.1) retrievals from the Atmospheric Carbon Observations from Space (ACOS) algorithm, includes significant improvements in the X CO2 data product compared to older OCO-2 data versions. This work compares the v11.1 X CO2 from OCO-2 against X CO2 estimates collected from a global ground-based network known as the Total Carbon Column Observing Network (TCCON), OCO-2's primary validation source. The OCO-2 project provides a version of the Level 2 data product, called “lite” files that include calibrated and bias-corrected XCO2 values, accessible together with all OCO-2 data products through the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). This work shows that OCO-2 X CO2 observations made between September 2014 and December 2023, after quality filtering and the application of an averaging kernel correction, agree well with coincident TCCON data for all OCO-2 observational modes of land (nadir, glint, target) and ocean (glint). The aggregated, bias-corrected, and quality-filtered absolute average bias values are less than or equal to 0.20 parts per million (ppm) globally for all OCO-2 observation modes, where the biases do not indicate a statistically significant time dependence. The land nadir/glint mode has the lowest bias value of −0.03 ± 0.85 ppm.

54 ENVIRONMENTAL SCIENCES↗

An overview of data tools for representing and managing building information and performance data

Building information modeling (BIM) has been widely adopted for representing and exchanging building data across disciplines during building design and construction. However, BIM's use in the building operation phase is limited. With the increasing deployment of low-cost sensors and meters, as well as affordable digital storage and computing technologies, growing volumes of data have been collected from buildings, their energy services systems, and occupants. Such data are crucial to help decision makers understand what, how, and when energy is consumed in buildings—a critical step to improving building performance for energy efficiency, demand flexibility, and resilience. However, practical analyses and use of the collected data are very limited due to various reasons, including poor data quality, ad-hoc representation of data, and lack of data science skills. To unlock value from building data, there is a strong need for a toolchain to curate and represent building information and performance data in common standardized terminologies and schemas, to enable interoperability between tools and applications. This study selected and reviewed 24 data tools based on common use cases of data across the building life cycle, from design to construction, commissioning, operation, and retrofits. The selected data tools are grouped into three categories: (1) data dictionary or terminology, (2) data ontology and schemas, and (3) data platforms. The data are grouped into ten typologies covering most types of data collected in buildings. This study resulted in five main findings: (1) most data representation tools can represent their intended data typologies well, such as Green Button for smart meter data and Brick schema for metadata of sensors in buildings and HVAC systems, but none of the tools cover all ten types of data; (2) there is a need for data schemas to represent the basis of design data and metadata of occupant data; (3) standard terminologies such as those defined in BEDES are only adopted in a few data tools; (4) integrating data across various stages in the building life cycle remains a challenge; and (5) most data tools were developed and maintained by different parties for different purposes, their flexibility and interoperability can be improved to support broader use cases. Finally, recommendations for future research on building data tools are provided for the data and buildings community based on the FAIR principles to make data Findable, Accessible, Interoperable, and Reusable.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hospice Landscape Report

In July 2018, CMS requested assistance from Oak Ridge National Laboratory (ORNL) to provide expert data science support aimed at developing algorithms for data mining of medical data for operational and payment purposes. The project is intended to be exploratory: work is aimed at alleviating challenges associated with improper payments, specifically audit methodologies and targeting and changes to risk scores. Project goals include developing new, sophisticated methods for audit targeting and improved profile– payment error correlations, specifically focused on Medicare Part C, the program under which MAOs provide health care services to beneficiaries. ORNL conducted RADV analyses against RAPS and EDS data as well as a hospice landscape analysis per a January 2015 dataset that included Medicare beneficiaries who were in hospice in 2017 and 2018. Ongoing work under this project also involves development of predictive models for RADV investigations and hospice landscape.

97 MATHEMATICS AND COMPUTING↗

Distributed Resources for the Earth System Grid Federation (ESGF) Advanced Management (DREAM). Final Report

Distributed Resources for the Earth System Grid Federation (ESGF) Advanced Management (DREAM) is a proposed system that will enable data from an infinite number of diverse sources to be organized and accessed from anywhere using any handheld or other computer device. The approach offers a powerful roadmap for the creation and integration of a unified knowledge base of an entire ecosystem, including its many geophysical, geographical, social, political, agricultural, energy, transportation, and cyber aspects. The resulting aggregation of data has the potential to generate an informational universe of unprecedented size that has never before been possible due to the prohibitive costs, managerial complexity, and technical barriers associated with ever-changing exponential-growth data flows. We envision that DREAM will accelerate discovery by enabling climate researchers, among other types of researchers, to manage, analyze, and visualize data from earth-scale measurements and simulations. DREAM’s success will be built on proven components that leverage existing services and resources. A key building block for DREAM will be the ESGF, chaired by Dean N. Williams. Expanding on the existing ESGF, the project will ensure that the access, storage, movement, and analysis of the large quantities of data that are processed and produced by diverse science projects can be dynamically distributed with proper resource management. Much of the Office of Science data is currently generated by multiple stand-alone facilities. DREAM can collect data accumulated from these facilities and incorporate it into a fully integrated network accessible from anywhere in the world. The result is a completely new paradigm shift for data management, analysis, and visualization enabling researchers to: Manage their calculations, data, tools, and research results; Ensure that all data are sharable, reproducible and (re)usable—accompanied by appropriate metadata describing its provenance, syntax, and semantics at creation; Advance application performance by selectively adapting APIs and services in response to scientific requirements and architectural complexities; and Provide scalable interactive resource management—navigate data and metadata at multiple levels, provide architecture-aware data integration, analysis and visualization tools. We will engage closely with DOE, NASA, and NOAA science groups working at the leading edge of computing. These engagements—in domains such as biology, climate, and hydrology—will allow us to advance disciplinary science goals and inform our development of technologies that can accelerate discovery across DOE more broadly. We will advertise and promote our technologies via dedicated workshops, tutorials, and sessions at conferences, stand-alone events with broad inter-disciplinary invitation, and engagements with leadership facilities.

54 ENVIRONMENTAL SCIENCES↗

Turbulence theories and statistical closure approaches

When discussing research in physics and in science more generally, it is common to ascribe equal importance to the three components of the scientific trinity: theoretical, experimental, and computational studies. This review will explore the future of modern turbulence theory by tracing its history, which began in earnest with Kolmogorov’s 1941 analysis of turbulence cascade and inertial range [A.N. Kolmogorov, Dokl. Akad. Nauk SSSR, 30, 299, (1941); 32, 19, (1941)]. The 80th Anniversary of Kolmogorov’s landmark study is a welcome opportunity to survey the achievements and evaluate the future of the theoretical approach of turbulence research. Over the years, turbulence theories have been critically important in laying the foundation of our understanding of the nature of turbulent flows. In particular, the Direct Interaction Approximation (DIA) [R.H. Kraichnan, J. Fluid Mech., 5, 497 (1959)] and its subsequent development, known as the statistical closure approach, can be identified as perhaps the most profound single advancement. The remarkable success of the statistical closure has furnished a platform to study such essential concepts as the energy transfer process and interacting scales, and the roles of the straining and sweeping motions. More recently, the quasi-Lagrangian formulation of V. L’vov & I. Procaccia and Kraichnan’s solvable passive scalar model provided powerful ways to explore another fundamental aspect of turbulent flows, the phenomena of intermittency, and the associated anomalous scaling exponents. In the meantime, the theory of fluid equilibria has been developed to describe the large-scale structures that can emerge from turbulent cascades of two-dimensional and geophysical flows at a later time. And yet, despite all these successes, analytical treatments suffer from mathematical complexities. As a result, the utility of theoretical approaches has been limited to relatively idealized flows. On the other hand, in recent decades, computational abilities and experimental facilities have reached an unprecedented scale. Looking beyond the horizon, the imminent deployment of exascale supercomputers will generate complete datasets of the entire flow field of key benchmark flows, allowing researchers to extract additional measurements concerning fully developed, complex turbulent flow fields far beyond those available from the statistical closure theories. Some other developments that could potentially influence the future course of turbulence theories include the advancement of machine learning, artificial intelligence, and data science; likely disruptions arising from the advent of quantum computation; and the increasingly prominent role of turbulence research in providing more accurate climate scientific data. Finally, turbulence theorists can leverage these developments by asking the right questions and developing advanced, sophisticated frameworks that will be able to predict and correlate vast amounts of data from the other two components of the trinity.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

First principles reactive simulation for equation of state prediction

The high cost of density functional theory (DFT) has hitherto limited the ab initio prediction of the equation of state (EOS). In this article, we employ a combination of large scale computing, advanced simulation techniques, and smart data science strategies to provide an unprecedented ab initio performance analysis of the high explosive pentaerythritol tetranitrate (PETN). Comparison to both experiment and thermochemical predictions reveals important quantitative limitations of DFT for EOS prediction and thus the assessment of high explosives. In particular, we find that DFT predicts the energy of PETN detonation products to be systematically too high relative to the unreacted neat crystalline material, resulting in an underprediction of the detonation velocity, pressure, and temperature at the Chapman–Jouguet state. The energetic bias can be partially accounted for by high-level electronic structure calculations of the product molecules. Furthermore we demonstrate a modeling strategy for mapping chemical composition across a wide parameter space with limited numerical data, the results of which suggest additional molecular species to consider in thermochemical modeling.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Recent advances and applications of deep learning methods in materials science

Deep learning (DL) is one of the fastest-growing topics in materials data science, with rapidly emerging applications spanning atomistic, image-based, spectral, and textual data modalities. DL allows analysis of unstructured data and automated identification of features. The recent development of large materials databases has fueled the application of DL methods in atomistic prediction in particular. In contrast, advances in image and spectral data have largely leveraged synthetic data enabled by high-quality forward models as well as by generative unsupervised DL methods. In this article, we present a high-level overview of deep learning methods followed by a detailed discussion of recent developments of deep learning in atomistic simulation, materials imaging, spectral analysis, and natural language processing. For each modality we discuss applications involving both theoretical and experimental data, typical modeling approaches with their strengths and limitations, and relevant publicly available software and datasets. We conclude the review with a discussion of recent cross-cutting work related to uncertainty quantification in this field and a brief perspective on limitations, challenges, and potential growth areas for DL methods in materials science.

36 MATERIALS SCIENCE↗

Data Analytics for Catalysis Predictions: Are We Ready Yet?

Catalysis informatics has received tremendous attention in recent years as a tool to design catalysts and discover unique descriptors that capture the relationships between chemical properties and catalytic performance. One of the stop-gaps in understanding catalytic effects, which is often ignored and limits the deployment of data science tools, relates to the lack of uniform data. The catalytic cleavage of C–X (X= H, C, N, and O) bonds is relevant to many fundamental catalytic processes. In this Perspective, we performed data analytics on four groups of C–X cleavage reactions that are common in production, upcycling, or reactive separation: the C–C cleavage in cyclopropyl alcohol, the C–H cleavage in hydroacylation reactions, the C–O cleavage in β-O-4 linkages, and the C–N cleavage in amides, using experimental data collected from the literature to understand their underlying correlations. Experimental variables of high impact are identified for each reaction by dimensionality reduction methods. We highlight the urgent need for experimental data sets that include full details on the reaction conditions, such as reagent concentration, reaction temperature, or time in machine-readable forms. We discuss the potential improvement of the data of these reactions and promising approaches such as autonomous experiments to fill the gaps in unbiased experimental data. Finally, we also address the early stage consideration of separation aspects in the experimental design of efficient catalytic systems for these fundamental examples of chemical reactivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

pyAMReX v23.08

The Python binding for AMReX, pyAMReX, bridges the worlds of block-structured codes and data science: it provides zero-copy application GPU data access for AI/ML, in situ analysis, application coupling and enables rapid, massively parallel prototyping.

Huebl, Axel↗