Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence (University of Delaware)

This report summarizes the activities, technical accomplishments, and outcomes of the RAPIDS2 Institute project at the University of Delaware (UD). The RAPIDS2 Institute was a large multi-institution project with the objective of assisting SciDAC and Office of Science application teams in the use of DOE supercomputing resources to achieve scientific breakthroughs. The UD team contributed to this effort through work on formal software verification. This thrust aims to reduce software developer time and effort, especially regarding debugging and testing, and to increase confidence in the correctness of the results computed by the software.

97 MATHEMATICS AND COMPUTING

Using data-science approaches to unravel insights for enhanced transport of lithium ions in single-ion conducting polymer electrolyte

Solid polymer electrolytes have yet to achieve the an ionic conductivity > 1 mS/cm at room temperature for realistic applications. This target implies the need to reduce the effective energy barriers of ion transport in polymer electrolytes to around 20 kJ/mol. In this work, we combine information extracted from existing experimental results with theoretical calculations to provide insights into ion transport in single-ion conductors (SICs) with a focus on lithium ion SICs. Through the analysis of temperature-dependent ionic conductivity data obtained from the literature, we evaluate different methods of extracting energy barriers for lithium transport. The traditional Arrhenius fit to the temperature-dependent ionic conductivity data indicates that the Meyer-Neldel rule holds for SICs. However, the values of the fitting parameters remain unphysical. Our modified approach based on recent work (Macromolecules, 56, 15, 6051(2023)), which incorporates a fixed pre-exponential factor, reveals that the energy barriers exhibit temperature dependence over a wide range of temperatures. Using this approach, we identify a series of anions leading to the energy barriers less than 30 kJ/mol, which include trifluoromethane sulfonimide (TFSI), fluoromethane sulfonimide (FSI), and boron-based organic anions. In our efforts to design the next generation of anions, which can exhibit the energy barriers less than 20 kJ/mol, we focused on boron-containing SICs, and performed density functional theory (DFT) based calculations to connect the chemical structures via the binding energy of cation (lithium)-anion pairs with the experimentally derived effective energy barriers for ion transport. Not only have we identified a correlation between the binding energy and the energy barriers, but we also propose a strategy to design new boron-based anions by using the correlation. This combined approach involving experiments and theoretical calculations is capable of facilitating the identification of promising new anions, which can exhibit ionic conductivity $> 1$ mS/cm near room temperature, thereby expediting the development of novel superionic single-ion conducting polymer electrolytes. The published datasets include all the temperature-dependent ionic conductivity collected from the literature with literature DOIs, DFT calculated binding energies, and python scripts to analyze data, construct statistical models, and generate plots.

36 MATERIALS SCIENCE

Using Data-Science Approaches to Unravel Insights for Enhanced Transport of Lithium Ions in Single-Ion Conducting Polymer Electrolytes

Solid polymer electrolytes have yet to achieve the desired ionic conductivity (>1 mS/cm) near room temperature required for many applications. This target implies the need to reduce the effective energy barriers for ion transport in polymer electrolytes to around 20 kJ/mol. In this work, we combine information extracted from existing experimental results with theoretical calculations to provide insights into ion transport in single-ion conductors (SICs) with a focus on lithium ion SICs. Through the analysis of temperature-dependent ionic conductivity data obtained from the literature, we evaluate different methods of extracting energy barriers for lithium transport. The traditional Arrhenius fit to the temperature-dependent ionic conductivity data indicates that the Meyer–Neldel rule holds for SICs. However, the values of the fitting parameters remain unphysical. Our modified approach based on recent work (Macromolecules 2023, 56, 15, 6051), which incorporates a fixed pre-exponential factor, reveals that the energy barriers exhibit temperature dependence over a wide range of temperatures. Using this approach, we identify anions leading to the energy barriers <30 kJ/mol, which include trifluoromethane sulfonimide (TFSI), fluoromethane sulfonimide (FSI), and boron-based organic anions. In our efforts to design the next generation of anions, which can exhibit the energy barriers <20 kJ/mol, we have performed density functional theory (DFT) based calculations to connect the chemical structures of boron-based anions via the binding energy of cation (lithium)-anion pairs with the experimentally derived effective energy barriers for ion hopping. Not only have we identified a correlation between the binding energy and the energy barriers, but we also propose a strategy to design new boron-based anions by using the correlation. This combined approach involving experiments and theoretical calculations is capable of facilitating the identification of promising new anions, which can exhibit ionic conductivity >1 mS/cm near room temperature, thereby expediting the development of novel superionic single-ion conducting polymer electrolytes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

The impact of urban configuration types on urban heat islands, air pollution, CO 2 emissions, and mortality in Europe: a data science approach

The world is becoming increasingly urbanized. As cities around the world continue to grow, it is important for urban planners and policymakers to understand how different urban configuration patterns affect the environment and human health. We aimed at identifying European urban configuration types, based on the Local Climate Zones categories and street design variables from Open Street Map, and evaluating their association with motorized traffic flows, Surface Urban Heat Island (SUHI) intensities, tropospheric nitrogen dioxide (NO 2 ), CO 2 per capita emissions and age-standardized mortality. We considered 946 European cities from 31 countries for the analysis defined in the 2018 Urban Audit database, of which 919 European cities were analysed. Data were collected at a 250 m × 250 m grid cell resolution. We divided all cities into five concentric rings based on the Burgess concentric urban planning model and calculated the mean values of all variables for each ring. First, to identify distinct urban configuration types, we applied the Uniform Manifold Approximation and Projection for Dimension Reduction method, followed by the k-means clustering algorithm. Next, statistical differences in exposures (including SUHI) and mortality between the resulting urban configuration types were evaluated using a Kruskal–Wallis test followed by a post-hoc Dunn's test. We identified four distinct urban configuration types characterising European cities: compact high density (n=246), open low-rise medium density (n=245), open low-rise low density (n=261), and green low density (n=167). Compact high density cities were a small size, had high population densities, and a low availability of natural areas. In contrast, green low-density cities were a large size, had low population densities, and a high availability of natural areas and cycleways. The open low-rise medium and low-density cities were a small to medium size with medium to low population densities and low to moderate availability of green areas. Motorised traffic flows and NO 2 exposure were significantly higher in compact high density and open low rise medium density cities when compared with green low density and open low-rise low density cities. Additionally, green low-density cities had a significantly lower SUHI effect compared with all other urban configuration types. Per person CO 2 emissions were significantly lower in compact high density cities compared with green low density cities. Lastly, green low density cities had significantly lower mortality rates when compared with all other urban configuration types. Our findings indicate that, although the compact city model is more sustainable, European compact cities still face challenges related to poor environmental quality and health. Our results have notable implications for urban and transport planning policies in Europe and contribute to the ongoing discussion on which city models can bring the greatest benefits for the environment, climate, and health.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING

Data Science in AD

Explore the source record for details and available documents.

Tang, Elaine [Fermilab]

Data for NB6 HBRR Science Design ORNL/TM-2025/3807

Data for the report (ORNL/TM-2025/3807) that describes the calculations and the Monte Carlo Ray Tracing simulations performed using the McStas package to determine the coatings and geometry for the NB-6 guide. It provides the information to inform the mechanical design, validation tests and verification that it meets the science requirements.

47 OTHER INSTRUMENTATION

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES

Data Science-Driven Discovery of Multimetallic Oxygen-cycle Electrocatalysts for Enhanced Energy Conversion

The overarching objective of this effort has been to combine state-of-the-art data science techniques, first principles analyses, and molecular-level characterization of electrocatalyst structure and reactivity to identify both in-situ mechanisms for degradation and transformation of electrocatalysts with highly complex catalytic structures and the impact of these transformations on catalytic activity. The primary catalysts of interest have been multielemental alloys, including high entropy alloys (HEA’s), which are characterized by a high degree of disorder and up to 20 different elements within a single nanoparticle. We have applied these strategies primarily to energy-critical oxygen cycle electrocatalytic reactions, including oxygen reduction (ORR), but we have also considered extensions to non-electrochemical chemistries such as ammonia synthesis and decomposition. We have made strong progress in the development of computational methods on both the level of machine learning methods development as well as first principles-based treatments of HEA’s, and we have leveraged these insights to propose promising HEA catalysts for the ORR. On the experimental side, we developed new HEA synthesis and characterization protocols relevant to these reactions and developed a database combining our experimental results with corresponding computational tools.

36 MATERIALS SCIENCE

Hands-On, Heads-Up: Blending Cyber T&E with Data Science-Driven Training in Jupyter Notebooks

In an era of increasingly sophisticated threats to critical infrastructure, cybersecurity professionals must be more than just aware; they must be immersed, agile, and equipped to operate in environments where failure is not an option. Nowhere is this truer than in the nuclear sector, where cyber-physical systems, regulatory scrutiny, and insider threat potential demand a new generation of hands-on, technically fluent defenders. This paper presents a unified training approach that integrates Cybersecurity Test and Evaluation (T&E) with data science techniques using Jupyter Notebooks as the interactive lab environment. The program centers on a modular, scenario-driven curriculum designed to build not just knowledge but practical capability in the assessment and defense of radiation detection systems, firmware interfaces, and operational security postures.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL

47 Tuc in Rubin Data Preview 1. Exploring Early LSST Data and Science Potential

We present analyses of the early data from Rubin Observatory’s Data Preview 1 (DP1) for the field of the globular cluster 47 Tuc. The DP1 data set for 47 Tuc includes four nights of observations from the Rubin Commissioning Camera (LSSTComCam), covering multiple bands (ugriy). We address challenges of crowding in the inner region of the cluster and toward the SMC in DP1, and demonstrate improved star–galaxy separation by fitting fifth-degree polynomials to the stellar loci in color–color diagrams and applying multidimensional sigma clipping. We compile a catalog of 3576 probable 47 Tuc member stars selected via a combination of isochrone, Gaia proper-motion, and color–color space matched filtering. We explore the sources of photometric scatter in the 47 Tuc color–color sequence, evaluating contributions from various potential sources, including differential extinction within the cluster. Finally, of the 72 well-characterized variables in the field, we recover three known variable stars, including two RR Lyrae and one eclipsing binary, in the coadd-based object catalog, and identify 62 in the difference image-based object catalog. Although the DP1 lightcurves have sparse temporal sampling, they appear to follow the patterns of densely sampled literature lightcurves well. Despite some data limitations for crowded-field stellar analysis, DP1 demonstrates the promising scientific potential for future LSST data releases.

Choi, Yumi [NSF National Optical-Infrared Astronom

Developing Fluorescence-Based Sensors to Support Rare Earth Element Separation

Rare earth elements (REEs) are essential to most renewable energy technologies. Unfortunately, as we transition to sustainable energy production, the demand for REEs is rapidly growing well beyond current rates of production. As a result, novel means of efficient, scalable, and easily adaptable methods for processing primary and recycle feedstocks are needed. Development and integration of sensors for highly selective in-line monitoring can support more efficient design and testing of such novel separation processes, as well as more cost-effective deployment of those separation flowsheets. Work here will explore the application of fluorescence spectroscopy, a highly sensitive and selective technique, to quantify multiple lanthanides in complex mixtures including known interferents or quenching agents. Results include identification of the optimal excitation wavelength and the limit of detection of various rare earth elements as well as the performance of data-science-based quantification approaches in streams where “unknowns” are present. Overall, the data science tools in conjunction with optical sensor data were able to quantify analytes in the presence of other lanthanides which can be anticipated in the actual industrial stream. Here we include characterization of lanthanides in a microfluidic device similar to those used in new process development. This study demonstrates the capability of utilizing fluorescence spectroscopy to quantify analytes in a complicated solution matrix, suggesting this is a successful approach for in-line monitoring to optimize the separation efficiency in an industrial stream.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato