Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data sciences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Data-driven materials research enabled by natural language processing and information extraction

Given the emergence of data science and machine learning throughout all aspects of society, but particularly in the scientific domain, there is increased importance placed on obtaining data. Data in materials science are particularly heterogeneous, based on the significant range in materials classes that are explored and the variety of materials properties that are of interest. This leads to data that range many orders of magnitude, and these data may manifest as numerical text or image-based information, which requires quantitative interpretation. The ability to automatically consume and codify the scientific literature across domains - enabled by techniques adapted from the field of natural language processing - therefore has immense potential to unlock and generate the rich datasets necessary for data science and machine learning. This review focuses on the progress and practices of natural language processing and text mining of materials science literature and highlights opportunities for extracting additional information beyond text contained in figures and tables in articles. Here, we discuss and provide examples for several reasons for the pursuit of natural language processing for materials, including data compilation, hypothesis development, and understanding the trends within and across fields. Current and emerging natural language processing methods along with their applications to materials science are detailed. We, then, discuss natural language processing and data challenges within the materials science domain where future directions may prove valuable.

36 MATERIALS SCIENCE↗

ThunderSecure: deploying real-time intrusion detection for 100G research networks by leveraging stream-based features and one-class classification network

Nowadays, data generated by large-scale scientific experiments are on the scale of petabytes per month. These data are transferred through dedicated high-bandwidth networks (40/100G) across distributed sites for processing, storage, and analysis. Like general purpose networks, research networks experience intrusions. However, monitoring anomalies in such high-speed network traffics is challenging given current cyber-infrastructure. Moreover, traditional network intrusion detection systems (NIDS) are signature based. However, anomaly patterns are difficult to define and that rulesets are often not updated frequently enough to reflect the changes of attack behaviors. We present ThunderSecure, a high-throughput, unsupervised learning-based intrusions detection system for 100G research networks. ThunderSecure implements an efficient packet processing and detection pipeline using multi-cores and GPUs. It extracts statistical and temporal features from real-time network data streams and feeds them to a one-class anomaly detection network. A baseline of normal distribution will be created based on the training observation. Testing traffic deviated from the learned profile will be marked as anomalies. We trained ThunderSecure on hundreds of billions of science data packets mirrored from two 100G network connections at Fermi National Accelerator Laboratory. The detection performance was evaluated on traffic captured from the same research network days and weeks after the training with different types of attack flows injected. Results show that ThunderSecure can recognize science data traffic captured long after the training and made nearly certain detection on the segment of the streams where anomalous flows were injected.

100G research network↗

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

PV Validation Hub

The Validation Hub will be a clearinghouse for the transfer of novel algorithms and software from the PV research community to industry. Potential algorithms tested in the Hub could include the estimation of various PV loss factors and the detection of various operational issues. The primary function of the Hub will be for developers to submit executable code which will run on hosted data sets. Developers will receive private reports on the accuracy and performance (e.g., run-time) of the submitted algorithms, and public high level summaries will be hosted. These summaries will indicate the organization who submitted the algorithm (e.g., links to GitHub pages, documentation websites, etc.), high-level accuracy metrics, and standardized performance metrics. These results will be stored in a publicly available database, accessible through the Hub, with the ability for users to sort and filter the results. In short, the Hub will be presented to public users as a collection of interactive leaderboards, organized around specific analysis tasks pertinent to the PV data science community. These tasks include things such as the estimation of various PV loss factors and the detection of various operational issues. We will present progress on the development of this hub, including preliminary results of comparative validation of PV data science algorithms and progress towards building the platform itself.

algorithm↗

Simulation and Emulation of X-Ray Diffraction from Dynamic Compression Experiments

Many important aspects of the dynamic thermo-mechanical response of materials occur at the mesoscale, i.e. a physical scale of interactions smaller than what can be adequately described by homogenous behaviors, yet larger than the scale of the atomic lattice. Concurrent advancements in computational power, continuum theory, and experimental diagnostics are enabling unprecedented understanding of such interactions. However, we cannot develop a sufficient level of confidence in such mesoscale capability until the constitutive description of the underlying constituents is reliably representative of their actual physical behavior. Therefore, there is a strong need to combine experimental, modeling, and data-science techniques to validate models of the thermomechanical response of individual single crystals. One experimental diagnostic with high potential impact to shock physics and materials science is in-situ x-ray diffraction. This paper is primarily focused on simulation of x-ray diffraction in shock physics, but with an aim toward quantifying parametric uncertainty of simulation models. Here, we develop and demonstrate a data-science and model-driven approach to constrain the parameterization of continuum models of crystal lattice deformation associated with the shock response of crystalline materials. The framework is built around the connection between continuum hydrodynamic simulations of lattice deformation and a new Bragg diffraction simulation code, BarberShop. The dynamic deformation of a crystal lattice is modeled using the DiscoFlux model within an arbitrary Lagrangian-Eulerian hydrodynamic code, FLAG. These detailed continuum simulations of lattice deformation can be computationally slow, thus a statistical model is used to emulate the evolution of lattice deformation fields in time and across the considered model parameter space. Emulated lattice deformation fields can then be generated rapidly for any combination of physics model parameters. In turn, these fields can be fed into BarberShop to realize a rapid prediction of Bragg diffraction patterns associated with particular values of physics model parameters. The framework enables parameterization of the single crystal model to obtain Bragg diffraction patterns that most closely resemble a corresponding measurement. Furthermore, the framework naturally provides sensitivities of the lattice deformation to the physics parameters. We highlight the utility of this framework through the application to a synthetic closed-loop inverse problem leading to the parameterization of a single crystal material model. As a model problem, we consider the dynamic response of the energetic molecular crystal, cyclotrimethylenetrinitramine (or RDX), under dynamic compression induced by simulated flyer plate impact experiments.

36 MATERIALS SCIENCE↗

The 2025 “Hacking Limnology” Workshop Series and DSOS Virtual Summit: A Half Decade of Data‐Intensive Aquatic Science

The 5th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) “Hacking Limnology” Workshop and 6th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 21–25 July 2025. As in previous years (Fig. 1; Meyer and Zwart 2020; Meyer et al. 2021b, 2021c, 2022, 2024), the virtual workshops and summit were free of charge, the content was formatted to allow for broad engagement from a globally distributed audience, and workshop materials and recordings were made available on the AEMON-J/DSOS archive (Meyer et al. 2021a). In contrast to previous years, which primarily focused on inland aquatic ecosystems, this year's workshops and summit showcased a notable plurality of ecosystem types, with workshops spanning marine, riverine, and lacustrine environments. The weeklong event brought together researchers and practitioners interested in the nexus of data science, open science, and the aquatic sciences, hosting between 47 and 65 attendees at a single time and a higher number of registrants (n = 389), who might opt to access the material asynchronously.

Meyer, Michael F. [US Geological Survey, Portland,↗

NGEE Arctic Authorship Guidelines

Authorship Guidelines were developed to help facilitate trust among team members as we span multiple institutions, scientific disciplines, and career stages. NGEE Arctic was built on a foundation of open science, data sharing, and collaboration. In Phase 4 of the project, it was particularly important to keep this foundation in mind as we develop new collaborations across the Arctic. Included in this package is one *.pdf. The Next-Generation Ecosystem Experiments in the Arctic (NGEE Arctic) project is a research effort to reduce uncertainty in the Department of Energy’s Energy Exascale Earth System Model (E3SM) by developing a predictive understanding of Arctic tundra ecosystems underlain by permafrost and to quantify feedbacks from the Arctic tundra to the Earth system. NGEE Arctic is supported by the Department of Energy's Office of Biological and Environmental Research. Over Phases 1–3, observations made by the NGEE Arctic team across a gradient of permafrost landscapes in Arctic Alaska improved the representation of tundra processes in the land surface component of E3SM (the E3SM Land Model, ELM). Model improvements emphasized unique aspects of permafrost environments and explored reductions in model complexity while retaining predictive power. The Arctic-informed ELM developed by NGEE Arctic has been used to make novel predictions on processes ranging from permafrost thaw to soil biogeochemical cycling to Earth system feedbacks associated with the unique characteristics of tundra plants. In Phase 4, the NGEE Arctic team is evaluating our new predictive understanding under novel conditions across the Arctic domain. In collaboration with partners at long-term pan-Arctic research sites we are examining whether an Arctic-informed ELM can faithfully simulate interactions among surface and subsurface processes at site, regional, and pan-Arctic scales. In turn, we are using variety of tools to dynamically extend and evaluate ELM inference, with an emphasis on data synthesis and pan-Arctic model evaluation, reintegration of code with an evolving E3SM, scaling across heterogeneous Arctic landscapes, and the appropriate representation of the impacts of increasingly frequent Arctic disturbances.

Iversen, Colleen [ORNL] (ORCID:0000000182933450)↗

Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats

Abstract Research can be more transparent and collaborative by using Findable, Accessible, Interoperable, and Reusable (FAIR) principles to publish Earth and environmental science data. Reporting formats—instructions, templates, and tools for consistently formatting data within a discipline—can help make data more accessible and reusable. However, the immense diversity of data types across Earth science disciplines makes development and adoption challenging. Here, we describe 11 community reporting formats for a diverse set of Earth science (meta)data including cross-domain metadata (dataset metadata, location metadata, sample metadata), file-formatting guidelines (file-level metadata, CSV files, terrestrial model data archiving), and domain-specific reporting formats for some biological, geochemical, and hydrological data (amplicon abundance tables, leaf-level gas exchange, soil respiration, water and sediment chemistry, sensor-based hydrologic measurements). More broadly, we provide guidelines that communities can use to create new (meta)data formats that integrate with their scientific workflows. Such reporting formats have the potential to accelerate scientific discovery and predictions by making it easier for data contributors to provide (meta)data that are more interoperable and reusable.

54 ENVIRONMENTAL SCIENCES↗

Optical versus radiographic imaging and tomography: introduction to the ROADS feature issue

Optical imaging is an ancient branch of imaging dating back to thousands of years. Radiographic imaging and tomography (RadIT), including the first use of X-rays by Wilhelm Röntgen, and then, $γ$ -rays, energetic charged particles, neutrons, etc. are about 130 years young. The synergies between optical and radiographic imaging can be cast in the framework of these building blocks: Physics, Sources, Detectors, Methods, and Data Science, as described in Appl. Opt. 61, RDS1 (2022). Optical imaging has expanded to include three-dimensional (3D) tomography (including holography), due in to part the invention of optical (including infrared) lasers. RadIT are intrinsically 3D because of the penetrating power of ionizing radiation. Both optical imaging and tomography (OIT) and RadIT are evolving into even higher dimensional regimes, such as time-resolved tomography (4D) and temporarily and spectroscopically resolved tomography (4D + ). Further advances in OIT and RadIT will continue to be driven by desires for higher information yield, higher resolutions, and higher probability models with reduced uncertainties. Synergies in quantum physics, laser-driven sources, low-cost detectors, data-driven methods, automated processing of data, and artificially intelligent data acquisition protocols will be beneficial to both branches of imaging in many applications. These topics, along with an overview of the Radiography, Applied Optics, and Data Science virtual feature issue, are discussed here.

47 OTHER INSTRUMENTATION↗

LSST Undergraduate Internships at Fermilab

The LSST Data Science for Undergraduate summer internship program focuses on data-driven astronomy for undergraduates at Fermilab’s Cosmic Physics Center. The internship activities focus on the development and implementation of data-driven investigatory techniques that will aid in LSST science, as well as prepare the undergraduates for future work in LSST. In particular, the internship activities include a number of opportunities for undergraduates to learn other skills critical for working on LSST --- data science research techniques, software development and engineering, science communication training.

79 ASTRONOMY AND ASTROPHYSICS↗

Toward Improved Regional Hydrological Model Performance Using State-Of-The-Science Data-Informed Soil Parameters

Accurate soil moisture and streamflow data are an aspirational need of many hydrologically relevant fields. Model simulated soil moisture and streamflow hold promise but models require validation prior to application. Calibration methods are commonly used to improve model fidelity but misrepresentation of the true dynamics remains a challenge. In this study, we leverage soil parameter estimates from the Soil Survey Geographic (SSURGO) database and the probability mapping of SSURGO (POLARIS) to improve the representation of hydrologic processes in the Weather Research and Forecasting Hydrological modeling system (WRF-Hydro) over a central California domain. Our results show WRF-Hydro soil moisture exhibits increased correlation coefficients ( r ), reduced biases, and increased Kling-Gupta Efficiencies (KGEs) across seven in situ soil moisture observing stations after updating the model's soil parameters according to POLARIS. Compared to four well-established soil moisture data sets including Soil Moisture Active Passive data and three Phase 2 North American Land Data Assimilation System land surface models, our POLARIS-adjusted WRF-Hydro simulations produce the highest mean KGE (0.69) across the seven stations. More importantly, WRF-Hydro streamflow fidelity also increases, especially in the case where the model domain is set up with SSURGO-informed total soil thickness. The magnitude and timing of peak flow events are better captured, r increases across nine United States Geological Survey stream gages, and the mean KGE across seven of the nine gages increases from 0.12 to 0.66. Our pre-calibration parameter estimate approach, which is transferable to other spatially distributed hydrological models, can substantially improve a model's performance, helping reduce calibration efforts and computational costs.

54 ENVIRONMENTAL SCIENCES↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Data from: “Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats”

This dataset contains supplementary information for a manuscript describing the ESS-DIVE (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) data repository's community data and metadata reporting formats. The purpose of creating the ESS-DIVE reporting formats was to provide guidelines for formatting some of the diverse data types that can be found in the ESS-DIVE repository. The 6 teams of community partners who developed the reporting formats included scientists and engineers from across the Department of Energy National Lab network. Additionally, during the development process, 247 individuals representing 128 institutions provided input on the formats. The primary files in this dataset are 10 data and metadata crosswalk for ESS-DIVE’s reporting formats (all files ending in _crosswalk.csv). The crosswalks compare elements used in each of the reporting formats to other related standards and data resources (e.g., repositories, datasets, data systems). This dataset also contains additional files recommended by ESS-DIVE’s file-level metadata reporting format. Each data file has an associated dictionary (files ending in _dd.csv) which provide a brief description of each standard or data resource consulted in the data reporting format development process. The flmd.csv file describes each file contained within the dataset.

54 ENVIRONMENTAL SCIENCES↗

Dani Sleight Intern Poster

The NRDS Portal is a login-based data storage solution and science data gateway for researchers to centralize and analyze data before publication. Users can upload, edit, review, and approve their own datasets within the site to eventually be published for public use on the main NRDS site. More development was needed to extend NRDS Portal with new Artificial Intelligence features.

99 - GENERAL AND MISCELLANEOUS↗

SIRIUS: Science-Driven Data Management for Multi-Tiered Storage

The data sets being generated by large applications on very large-scale systems are increasing in both size and complexity. At the same time, there are new ways available to store and access these data sets. The goal in this project is to develop software that applications can use to make use of new and existing storage technologies in more sophisticated ways. One challenge in scientific data management is handling ‘hot’ vs ‘cold’ data. Data that is hot is data that is needed (or will be needed soon) in order for the program to continue progressing, while cold data is either output (and so will not be need further during the life of the program) or will not be needed until significantly later in the program’s run. Hot data should be stored in a way that allows fast access. On most systems, economic factors lead to an inverse relationship between storage performance and storage capacity and so fast access storage is limited. This makes it important to correctly place hot and cold data and avoid cold data unnecessarily consuming precious resources. In this reporting period, we addressed this challenge in various ways and at various levels. Data management frameworks offer only limited control to applications in how data is stored. We have added software capabilities for seamlessly moving data between layers of the storage technology using promote and demote functions to existing software frameworks. This gives direct control to applications in deciding what priority data receives. Additionally, we integrated different storage layer management frameworks in order to allow data to be exchanged and moved between storage layers in a consistent way across the application. Further, applications are not always able to directly decide what storage level makes sense for a given piece of data without an understanding of the underlying storage technologies. Data storage frameworks are often positioned to make these sorts of decisions in service of the application. We have added machine-learning based capabilities to data staging frameworks in order to make intelligent decisions about where data should be stored given learning about patterns in previous usage of similar data.

97 MATHEMATICS AND COMPUTING↗

Carbon Storage Technical Viability Approach (CS TVA): An Integrated Approach for Feasibility and Data Resource Assessment

There is currently a poor understanding and lack of workflow to understand the technical viability of carbon storage spatially. To address this gap, the multi-faceted Carbon Storage Technical Viability Approach (CS TVA) is being developed to incorporate CO2 storage resources, environmental and socio-economic justice (EJ/SJ) factors to enable more comprehensive assessments. The CS TVA includes a (1) matrix framework, (2) an integrated and labeled database, (3) a data availability assessment workflow, and (4) spatial data availability assessment results. This approach leverages spatial and data science analytics to communicate data density, uncertainty, and gaps. The workflow can be applied in whole or in part, based on user needs.

Rodriguez, Neyda Cordero↗

The 2024 “Hacking Limnology” Workshop Series and Virtual Summit: Increasing Inclusion, Participation, and Representation in the Aquatic Sciences

The 4th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) Hacking Limnology Workshop and 5th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 15–19 July 2024. During the week, these joint communities engaged in activities at the intersection of big data, open science, modeling, remote sensing, and the aquatic sciences. The weeklong event, with over 100 aquatic science practitioners and enthusiasts, followed a similar structure to previous years, comprising three days of workshops followed by two days of the virtual summit.

54 ENVIRONMENTAL SCIENCES↗