Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data sciences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Comparisons of the v11.1 Orbiting Carbon Observatory‐2 (OCO‐2) X CO2 Measurements With GGG2020 TCCON

The Orbiting Carbon Observatory 2 (OCO-2) is NASA's first Earth observation satellite mission dedicated to studying the sources and sinks of carbon dioxide (CO 2 ) on a global scale. The observations of reflected sunlight are inverted in a retrieval algorithm to produce estimates of the dry air mole-fractions of CO 2 (X CO2 ). The OCO-2 Level 2 data release, version 11.1 (v11.1) retrievals from the Atmospheric Carbon Observations from Space (ACOS) algorithm, includes significant improvements in the X CO2 data product compared to older OCO-2 data versions. This work compares the v11.1 X CO2 from OCO-2 against X CO2 estimates collected from a global ground-based network known as the Total Carbon Column Observing Network (TCCON), OCO-2's primary validation source. The OCO-2 project provides a version of the Level 2 data product, called “lite” files that include calibrated and bias-corrected XCO2 values, accessible together with all OCO-2 data products through the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). This work shows that OCO-2 X CO2 observations made between September 2014 and December 2023, after quality filtering and the application of an averaging kernel correction, agree well with coincident TCCON data for all OCO-2 observational modes of land (nadir, glint, target) and ocean (glint). The aggregated, bias-corrected, and quality-filtered absolute average bias values are less than or equal to 0.20 parts per million (ppm) globally for all OCO-2 observation modes, where the biases do not indicate a statistically significant time dependence. The land nadir/glint mode has the lowest bias value of −0.03 ± 0.85 ppm.

54 ENVIRONMENTAL SCIENCES↗

An overview of data tools for representing and managing building information and performance data

Building information modeling (BIM) has been widely adopted for representing and exchanging building data across disciplines during building design and construction. However, BIM's use in the building operation phase is limited. With the increasing deployment of low-cost sensors and meters, as well as affordable digital storage and computing technologies, growing volumes of data have been collected from buildings, their energy services systems, and occupants. Such data are crucial to help decision makers understand what, how, and when energy is consumed in buildings—a critical step to improving building performance for energy efficiency, demand flexibility, and resilience. However, practical analyses and use of the collected data are very limited due to various reasons, including poor data quality, ad-hoc representation of data, and lack of data science skills. To unlock value from building data, there is a strong need for a toolchain to curate and represent building information and performance data in common standardized terminologies and schemas, to enable interoperability between tools and applications. This study selected and reviewed 24 data tools based on common use cases of data across the building life cycle, from design to construction, commissioning, operation, and retrofits. The selected data tools are grouped into three categories: (1) data dictionary or terminology, (2) data ontology and schemas, and (3) data platforms. The data are grouped into ten typologies covering most types of data collected in buildings. This study resulted in five main findings: (1) most data representation tools can represent their intended data typologies well, such as Green Button for smart meter data and Brick schema for metadata of sensors in buildings and HVAC systems, but none of the tools cover all ten types of data; (2) there is a need for data schemas to represent the basis of design data and metadata of occupant data; (3) standard terminologies such as those defined in BEDES are only adopted in a few data tools; (4) integrating data across various stages in the building life cycle remains a challenge; and (5) most data tools were developed and maintained by different parties for different purposes, their flexibility and interoperability can be improved to support broader use cases. Finally, recommendations for future research on building data tools are provided for the data and buildings community based on the FAIR principles to make data Findable, Accessible, Interoperable, and Reusable.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hospice Landscape Report

In July 2018, CMS requested assistance from Oak Ridge National Laboratory (ORNL) to provide expert data science support aimed at developing algorithms for data mining of medical data for operational and payment purposes. The project is intended to be exploratory: work is aimed at alleviating challenges associated with improper payments, specifically audit methodologies and targeting and changes to risk scores. Project goals include developing new, sophisticated methods for audit targeting and improved profile– payment error correlations, specifically focused on Medicare Part C, the program under which MAOs provide health care services to beneficiaries. ORNL conducted RADV analyses against RAPS and EDS data as well as a hospice landscape analysis per a January 2015 dataset that included Medicare beneficiaries who were in hospice in 2017 and 2018. Ongoing work under this project also involves development of predictive models for RADV investigations and hospice landscape.

97 MATHEMATICS AND COMPUTING↗

Distributed Resources for the Earth System Grid Federation (ESGF) Advanced Management (DREAM). Final Report

Distributed Resources for the Earth System Grid Federation (ESGF) Advanced Management (DREAM) is a proposed system that will enable data from an infinite number of diverse sources to be organized and accessed from anywhere using any handheld or other computer device. The approach offers a powerful roadmap for the creation and integration of a unified knowledge base of an entire ecosystem, including its many geophysical, geographical, social, political, agricultural, energy, transportation, and cyber aspects. The resulting aggregation of data has the potential to generate an informational universe of unprecedented size that has never before been possible due to the prohibitive costs, managerial complexity, and technical barriers associated with ever-changing exponential-growth data flows. We envision that DREAM will accelerate discovery by enabling climate researchers, among other types of researchers, to manage, analyze, and visualize data from earth-scale measurements and simulations. DREAM’s success will be built on proven components that leverage existing services and resources. A key building block for DREAM will be the ESGF, chaired by Dean N. Williams. Expanding on the existing ESGF, the project will ensure that the access, storage, movement, and analysis of the large quantities of data that are processed and produced by diverse science projects can be dynamically distributed with proper resource management. Much of the Office of Science data is currently generated by multiple stand-alone facilities. DREAM can collect data accumulated from these facilities and incorporate it into a fully integrated network accessible from anywhere in the world. The result is a completely new paradigm shift for data management, analysis, and visualization enabling researchers to: Manage their calculations, data, tools, and research results; Ensure that all data are sharable, reproducible and (re)usable—accompanied by appropriate metadata describing its provenance, syntax, and semantics at creation; Advance application performance by selectively adapting APIs and services in response to scientific requirements and architectural complexities; and Provide scalable interactive resource management—navigate data and metadata at multiple levels, provide architecture-aware data integration, analysis and visualization tools. We will engage closely with DOE, NASA, and NOAA science groups working at the leading edge of computing. These engagements—in domains such as biology, climate, and hydrology—will allow us to advance disciplinary science goals and inform our development of technologies that can accelerate discovery across DOE more broadly. We will advertise and promote our technologies via dedicated workshops, tutorials, and sessions at conferences, stand-alone events with broad inter-disciplinary invitation, and engagements with leadership facilities.

54 ENVIRONMENTAL SCIENCES↗

Turbulence theories and statistical closure approaches

When discussing research in physics and in science more generally, it is common to ascribe equal importance to the three components of the scientific trinity: theoretical, experimental, and computational studies. This review will explore the future of modern turbulence theory by tracing its history, which began in earnest with Kolmogorov’s 1941 analysis of turbulence cascade and inertial range [A.N. Kolmogorov, Dokl. Akad. Nauk SSSR, 30, 299, (1941); 32, 19, (1941)]. The 80th Anniversary of Kolmogorov’s landmark study is a welcome opportunity to survey the achievements and evaluate the future of the theoretical approach of turbulence research. Over the years, turbulence theories have been critically important in laying the foundation of our understanding of the nature of turbulent flows. In particular, the Direct Interaction Approximation (DIA) [R.H. Kraichnan, J. Fluid Mech., 5, 497 (1959)] and its subsequent development, known as the statistical closure approach, can be identified as perhaps the most profound single advancement. The remarkable success of the statistical closure has furnished a platform to study such essential concepts as the energy transfer process and interacting scales, and the roles of the straining and sweeping motions. More recently, the quasi-Lagrangian formulation of V. L’vov & I. Procaccia and Kraichnan’s solvable passive scalar model provided powerful ways to explore another fundamental aspect of turbulent flows, the phenomena of intermittency, and the associated anomalous scaling exponents. In the meantime, the theory of fluid equilibria has been developed to describe the large-scale structures that can emerge from turbulent cascades of two-dimensional and geophysical flows at a later time. And yet, despite all these successes, analytical treatments suffer from mathematical complexities. As a result, the utility of theoretical approaches has been limited to relatively idealized flows. On the other hand, in recent decades, computational abilities and experimental facilities have reached an unprecedented scale. Looking beyond the horizon, the imminent deployment of exascale supercomputers will generate complete datasets of the entire flow field of key benchmark flows, allowing researchers to extract additional measurements concerning fully developed, complex turbulent flow fields far beyond those available from the statistical closure theories. Some other developments that could potentially influence the future course of turbulence theories include the advancement of machine learning, artificial intelligence, and data science; likely disruptions arising from the advent of quantum computation; and the increasingly prominent role of turbulence research in providing more accurate climate scientific data. Finally, turbulence theorists can leverage these developments by asking the right questions and developing advanced, sophisticated frameworks that will be able to predict and correlate vast amounts of data from the other two components of the trinity.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

First principles reactive simulation for equation of state prediction

The high cost of density functional theory (DFT) has hitherto limited the ab initio prediction of the equation of state (EOS). In this article, we employ a combination of large scale computing, advanced simulation techniques, and smart data science strategies to provide an unprecedented ab initio performance analysis of the high explosive pentaerythritol tetranitrate (PETN). Comparison to both experiment and thermochemical predictions reveals important quantitative limitations of DFT for EOS prediction and thus the assessment of high explosives. In particular, we find that DFT predicts the energy of PETN detonation products to be systematically too high relative to the unreacted neat crystalline material, resulting in an underprediction of the detonation velocity, pressure, and temperature at the Chapman–Jouguet state. The energetic bias can be partially accounted for by high-level electronic structure calculations of the product molecules. Furthermore we demonstrate a modeling strategy for mapping chemical composition across a wide parameter space with limited numerical data, the results of which suggest additional molecular species to consider in thermochemical modeling.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Recent advances and applications of deep learning methods in materials science

Deep learning (DL) is one of the fastest-growing topics in materials data science, with rapidly emerging applications spanning atomistic, image-based, spectral, and textual data modalities. DL allows analysis of unstructured data and automated identification of features. The recent development of large materials databases has fueled the application of DL methods in atomistic prediction in particular. In contrast, advances in image and spectral data have largely leveraged synthetic data enabled by high-quality forward models as well as by generative unsupervised DL methods. In this article, we present a high-level overview of deep learning methods followed by a detailed discussion of recent developments of deep learning in atomistic simulation, materials imaging, spectral analysis, and natural language processing. For each modality we discuss applications involving both theoretical and experimental data, typical modeling approaches with their strengths and limitations, and relevant publicly available software and datasets. We conclude the review with a discussion of recent cross-cutting work related to uncertainty quantification in this field and a brief perspective on limitations, challenges, and potential growth areas for DL methods in materials science.

36 MATERIALS SCIENCE↗

Data Analytics for Catalysis Predictions: Are We Ready Yet?

Catalysis informatics has received tremendous attention in recent years as a tool to design catalysts and discover unique descriptors that capture the relationships between chemical properties and catalytic performance. One of the stop-gaps in understanding catalytic effects, which is often ignored and limits the deployment of data science tools, relates to the lack of uniform data. The catalytic cleavage of C–X (X= H, C, N, and O) bonds is relevant to many fundamental catalytic processes. In this Perspective, we performed data analytics on four groups of C–X cleavage reactions that are common in production, upcycling, or reactive separation: the C–C cleavage in cyclopropyl alcohol, the C–H cleavage in hydroacylation reactions, the C–O cleavage in β-O-4 linkages, and the C–N cleavage in amides, using experimental data collected from the literature to understand their underlying correlations. Experimental variables of high impact are identified for each reaction by dimensionality reduction methods. We highlight the urgent need for experimental data sets that include full details on the reaction conditions, such as reagent concentration, reaction temperature, or time in machine-readable forms. We discuss the potential improvement of the data of these reactions and promising approaches such as autonomous experiments to fill the gaps in unbiased experimental data. Finally, we also address the early stage consideration of separation aspects in the experimental design of efficient catalytic systems for these fundamental examples of chemical reactivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Generalized Canonical Polyadic Tensor Decomposition

Tensor decomposition is a fundamental unsupervised machine learning method in data science, with applications including network analysis and sensor data processing. This work develops a generalized canonical polyadic (GCP) low-rank tensor decomposition that allows other loss functions besides squared error. For instance, we can use logistic loss or Kullback--Leibler divergence, enabling tensor decomposition for binary or count data. We present a variety of statistically motivated loss functions for various scenarios. We provide a generalized framework for computing gradients and handling missing data that enables the use of standard optimization methods for fitting the model. Furthermore, we demonstrate the flexibility of the GCP decomposition on several real-world examples including interactions in a social network, neural activity in a mouse, and monthly rainfall measurements in India.

97 MATHEMATICS AND COMPUTING↗

pyAMReX v23.08

The Python binding for AMReX, pyAMReX, bridges the worlds of block-structured codes and data science: it provides zero-copy application GPU data access for AI/ML, in situ analysis, application coupling and enables rapid, massively parallel prototyping.

Huebl, Axel↗

Roadmap on data-centric materials science

Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

36 MATERIALS SCIENCE↗

A Modular System for Increasing Predictiveness for Extreme Climate Predictions

We know that climate change is poised to reshape our world, but we lack clear enough predictions about precisely how. The preponderance of these changes is associated with human activity, specifically the emission of CO 2 and other greenhouse gases. Problematically, projections of climate change continue to be marred by unacceptably large uncertainties which hamper informed decision-making and cost society a chance to adapt proactively and effectively. These uncertainties stem from deficiencies in predictions of future greenhouse gas emissions, but also from inaccuracies in the representation of the physical models used to predict the climate response to such emissions. The uncertainties in projections associated with the inaccurate representation of climate physics, chemistry and biology are similar to those that plagued the first global climate models developed fifty years ago, despite more than a factor 10 8 increase in computer performance. Our transformational question is then, how can the accuracy of climate projections be dramatically improved by applying recent advances in the computational and data sciences to train the models with the wealth of data being constantly collected about the ongoing changes in the climate system?

54 ENVIRONMENTAL SCIENCES↗

ZENN: A thermodynamics-inspired computational framework for heterogeneous data–driven modeling

Traditional entropy-based methods—such as cross-entropy loss in classification problems—have long been essential tools for representing the information uncertainty and physical disorder in data and for developing artificial intelligence algorithms. However, the rapid growth of data across various domains has introduced new challenges, particularly the integration of heterogeneous datasets with intrinsic disparities. To address this, we introduce a zentropy-enhanced neural network (ZENN), extending zentropy theory into the data science domain via intrinsic entropy, enabling more effective learning from heterogeneous data sources. ZENN simultaneously learns both energy and intrinsic entropy components, capturing the underlying structure of multisource data. To support this, we redesign the neural network architecture to better reflect the intrinsic properties and variability inherent in diverse datasets. We demonstrate the effectiveness of ZENN on classification tasks and energy landscape reconstructions, showing its superior generalization capabilities and robustness-particularly in predicting high-order derivatives. In image and text classification tasks, ZENN demonstrates superior generalization by introducing a learnable temperature variable that models latent multisource heterogeneity, allowing it to surpass state-of-the-art models on CIFAR-10/100, BBC News, and AG News. As a practical application in materials science, we employ ZENN to reconstruct the Helmholtz energy landscape of Fe3Pt using data generated from density functional theory and capture key material behaviors, including negative thermal expansion and the critical point in the temperature–pressure space. Overall, this work presents a zentropy-grounded framework for data-driven machine learning, positioning ZENN as a versatile and robust approach for scientific problems involving complex, heterogeneous datasets.

36 MATERIALS SCIENCE↗

Reviews and syntheses: The promise of big diverse soil data, moving current practices towards future potential

Abstract. In the age of big data, soil data are more available and richer than ever, but – outside of a few large soil survey resources – they remain largely unusable for informing soil management and understanding Earth system processes beyond the original study. Data science has promised a fully reusable research pipeline where data from past studies are used to contextualize new findings and reanalyzed for new insight. Yet synthesis projects encounter challenges at all steps of the data reuse pipeline, including unavailable data, labor-intensive transcription of datasets, incomplete metadata, and a lack of communication between collaborators. Here, using insights from a diversity of soil, data, and climate scientists, we summarize current practices in soil data synthesis across all stages of database creation: availability, input, harmonization, curation, and publication. We then suggest new soil-focused semantic tools to improve existing data pipelines, such as ontologies, vocabulary lists, and community practices. Our goal is to provide the soil data community with an overview of current practices in soil data and where we need to go to fully leverage big data to solve soil problems in the next century.

54 ENVIRONMENTAL SCIENCES↗

Integrated parameter and process learning for hydrologic and biogeochemical modules in Earth System Models

Focus area: Primary focal area #2; secondary focal area #3: Learning about parameters and processes of land surface hydrologic and biogeochemical models in Earth System models by integrating machine learning, physics, and big data. Science challenges: How do we maximally leverage big-data observations to improve hydrobiogeochemical process description and parameterization so that such modules more realistically capture hydrologic and vegetation responses and feedbacks under the future climate? For example, how can we leverage physics, limited observations of vegetation and streamflow to better estimate evapotranspiration, and, relatedly, net primary productivity, especially for drought areas? Vegetation plays a critical role in regional and global water cycles; however, existing vegetation models have failed to predict vegetation response to droughts (McDowell & Xu, 2017) , arctic greening (Keenan & Riley, 2018) , and critical transitions between forest and savanna (Hirota et al., 2011) . These studies suggest that when we build process-based models (PBM) parameterized from regional and global plant traits, we tend to poorly describe plant adaptation and local-scale competition processes. The models and their associated parameters assigned for different regions in the world are not capturing essential heterogeneity in vegetation responses at finer spatial scales. Many parameters of the land surface models control hydrology and vegetation dynamics at the same time. The heterogeneity in vegetation response is a function of (i) plant type, (ii) plant size, (iii) competition and succession, (iv) environmental controls, and (v) local variations due to the unique ecological community that are very difficult to describe (e.g., the size of gaps resulting from fire that facilitated the coexistence of pioneering species). In the demographic models, only factors (i) and (iv) were captured, and plant types were generally described only by leaf phenology and climate zones. With current demographic models, we generally consider more traits to define plant types (i) and calibrate these traits to consider factors (ii), (iii) and (iv); however, it is substantially challenging to scale to regional and global simulations due to trait variations across space (Ali et al., 2016). Moreover, it has been noted that hillslope processes, including ridge-to-valley flow and sunny vs. shady slopes are primary organizers of water, energy, and vegetation (Clark et al., 2015; Fan et al., 2019) . Although gradual improvements in the hydrologic model component in earth system models may reduce this error (at a remarkably slow pace), the long-term, gradual impact of hydrology on plant traits are not well captured. Recent work showed that the hydrologic controls exerted by groundwater and lateral flow are primary regulators of rooting depth (Fan et al., 2017) . Such hydrologic controls have seldom been reflected in vegetation model parameterizations.

54 ENVIRONMENTAL SCIENCES↗

Crowdsourcing the Frontier: Advancing Hybrid Physics‐ML Climate Simulation via a $\$$50,000 Kaggle Competition

Subgrid machine-learning (machine learning [ML]) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and ML researchers opened up the offline aspect of this problem to the broader ML and data science community with the release of ClimSim, a NeurIPS Data sets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art results on certain metrics such as zonal mean bias patterns and global Root Mean Squared Error, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

Environmental sciences↗

Measuring stellar populations, dust attenuation and ionized gas at kpc scales in 10010 nearby galaxies using the integral field spectroscopy from MaNGA

As one of the three major experiments of the fourth-generation Sloan Digital Sky Survey (SDSS-IV), the Mapping Nearby Galaxies at Apatch Point Observatory (MaNGA) survey has obtained high-quality integral field spectroscopy (IFS) with a resolution of 1–2 kpc for ∼ 10 4 galaxies in the local universe during its six-year operation from July 2014 through August 2020. It is crucial to reliably measure the physical properties of the different components in each spectrum before one can use the data for any scientific study. In the past years we have made lots of efforts to develop a novel technique of full spectral fitting, which estimates a model-independent dust attenuation curve from each spectrum, thus allowing us to break the degeneracy between dust attenuation and stellar population properties when fitting the spectrum with stellar population synthesis models. We have applied our technique to the final data release of MaNGA, and obtained measurements of stellar population properties and emission line parameters, as well as the kinematics and dust attenuation of both stellar and ionized gas components. In this paper we describe our technique and the content and format of our data products. The whole dataset is publicly available in Science Data Bank with the link https://doi.org/10.57760/sciencedb.j00113.00088 .

Physics↗