Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Science Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING

Data Science in AD

Explore the source record for details and available documents.

Tang, Elaine [Fermilab]

Data for NB6 HBRR Science Design ORNL/TM-2025/3807

Data for the report (ORNL/TM-2025/3807) that describes the calculations and the Monte Carlo Ray Tracing simulations performed using the McStas package to determine the coatings and geometry for the NB-6 guide. It provides the information to inform the mechanical design, validation tests and verification that it meets the science requirements.

47 OTHER INSTRUMENTATION

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES

Data Science-Driven Discovery of Multimetallic Oxygen-cycle Electrocatalysts for Enhanced Energy Conversion

The overarching objective of this effort has been to combine state-of-the-art data science techniques, first principles analyses, and molecular-level characterization of electrocatalyst structure and reactivity to identify both in-situ mechanisms for degradation and transformation of electrocatalysts with highly complex catalytic structures and the impact of these transformations on catalytic activity. The primary catalysts of interest have been multielemental alloys, including high entropy alloys (HEA’s), which are characterized by a high degree of disorder and up to 20 different elements within a single nanoparticle. We have applied these strategies primarily to energy-critical oxygen cycle electrocatalytic reactions, including oxygen reduction (ORR), but we have also considered extensions to non-electrochemical chemistries such as ammonia synthesis and decomposition. We have made strong progress in the development of computational methods on both the level of machine learning methods development as well as first principles-based treatments of HEA’s, and we have leveraged these insights to propose promising HEA catalysts for the ORR. On the experimental side, we developed new HEA synthesis and characterization protocols relevant to these reactions and developed a database combining our experimental results with corresponding computational tools.

36 MATERIALS SCIENCE

Hands-On, Heads-Up: Blending Cyber T&E with Data Science-Driven Training in Jupyter Notebooks

In an era of increasingly sophisticated threats to critical infrastructure, cybersecurity professionals must be more than just aware; they must be immersed, agile, and equipped to operate in environments where failure is not an option. Nowhere is this truer than in the nuclear sector, where cyber-physical systems, regulatory scrutiny, and insider threat potential demand a new generation of hands-on, technically fluent defenders. This paper presents a unified training approach that integrates Cybersecurity Test and Evaluation (T&E) with data science techniques using Jupyter Notebooks as the interactive lab environment. The program centers on a modular, scenario-driven curriculum designed to build not just knowledge but practical capability in the assessment and defense of radiation detection systems, firmware interfaces, and operational security postures.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL

47 Tuc in Rubin Data Preview 1. Exploring Early LSST Data and Science Potential

We present analyses of the early data from Rubin Observatory’s Data Preview 1 (DP1) for the field of the globular cluster 47 Tuc. The DP1 data set for 47 Tuc includes four nights of observations from the Rubin Commissioning Camera (LSSTComCam), covering multiple bands (ugriy). We address challenges of crowding in the inner region of the cluster and toward the SMC in DP1, and demonstrate improved star–galaxy separation by fitting fifth-degree polynomials to the stellar loci in color–color diagrams and applying multidimensional sigma clipping. We compile a catalog of 3576 probable 47 Tuc member stars selected via a combination of isochrone, Gaia proper-motion, and color–color space matched filtering. We explore the sources of photometric scatter in the 47 Tuc color–color sequence, evaluating contributions from various potential sources, including differential extinction within the cluster. Finally, of the 72 well-characterized variables in the field, we recover three known variable stars, including two RR Lyrae and one eclipsing binary, in the coadd-based object catalog, and identify 62 in the difference image-based object catalog. Although the DP1 lightcurves have sparse temporal sampling, they appear to follow the patterns of densely sampled literature lightcurves well. Despite some data limitations for crowded-field stellar analysis, DP1 demonstrates the promising scientific potential for future LSST data releases.

Choi, Yumi [NSF National Optical-Infrared Astronom

Developing Fluorescence-Based Sensors to Support Rare Earth Element Separation

Rare earth elements (REEs) are essential to most renewable energy technologies. Unfortunately, as we transition to sustainable energy production, the demand for REEs is rapidly growing well beyond current rates of production. As a result, novel means of efficient, scalable, and easily adaptable methods for processing primary and recycle feedstocks are needed. Development and integration of sensors for highly selective in-line monitoring can support more efficient design and testing of such novel separation processes, as well as more cost-effective deployment of those separation flowsheets. Work here will explore the application of fluorescence spectroscopy, a highly sensitive and selective technique, to quantify multiple lanthanides in complex mixtures including known interferents or quenching agents. Results include identification of the optimal excitation wavelength and the limit of detection of various rare earth elements as well as the performance of data-science-based quantification approaches in streams where “unknowns” are present. Overall, the data science tools in conjunction with optical sensor data were able to quantify analytes in the presence of other lanthanides which can be anticipated in the actual industrial stream. Here we include characterization of lanthanides in a microfluidic device similar to those used in new process development. This study demonstrates the capability of utilizing fluorescence spectroscopy to quantify analytes in a complicated solution matrix, suggesting this is a successful approach for in-line monitoring to optimize the separation efficiency in an industrial stream.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)

Enhancing risk and crisis communication with computational methods: A systematic literature review

Abstract Recent developments in risk and crisis communication (RCC) research combine social science theory and data science tools to construct effective risk messages efficiently. However, current systematic literature reviews (SLRs) on RCC primarily focus on computationally assessing message efficacy as opposed to message efficiency. We conduct an SLR to highlight any current computational methods that improve message construction efficacy and efficiency. We found that most RCC research focuses on using theoretical frameworks and computational methods to analyze or classify message elements that improve efficacy. For improving message efficiency, computational and manual methods are only used in message classification. Specifying the computational methods used in message construction is sparse. We recommend that future RCC research apply computational methods toward improving efficacy and efficiency in message construction. By improving message construction efficacy and efficiency, RCC messaging would quickly warn and better inform affected communities impacted by current hazards. Such messaging has the potential to save as many lives as possible.

Mathematical Methods In Social Sciences

Towards a Robust Adaptive Digital Twin for Fusion Applications

The development of a digital twin system for fusion applications is essential for enhancing the prediction, analysis, and optimization of complex plasma processes. Machine learning (ML), particularly deep learning has demonstrated strong capabilities in modeling such highly nonlinear and intricate systems. However, two critical challenges limit the deployment of deep learning-based digital twins: Uncertainty Quantification (UQ) and data drift. UQ is vital for ensuring trustworthy predictions, especially in decision-support scenarios. Additionally, data-driven models are often sensitive to changes in the underlying data distribution, such as shot-to-shot variations in fusion experiments, which can lead to performance degradation over time. To address these challenges, we are developing an uncertainty-aware, adaptive digital twin framework. Our approach incorporates deep learning models enhanced with Gaussian Process approximations for predictive uncertainty estimation, coupled with an online learning mechanism that enables continuous model adaptation to new experimental data. This adaptive capability allows the data driven models to respond effectively to evolving plasma behaviors and equipment conditions. Specifically, to mitigate the effects of shot-to-shot drift, our system updates itself incrementally as new data becomes available, improving both robustness and fidelity. Our vision is to evolve this data driven model into a self-sustaining digital twin system that leverages UQ based feedback to continuously refine itself and potentially support real-time decision making. This presentation will cover a brief background on uncertainty quantification for ML, our ongoing effort on development of UQ capabilities for ML, our data science pipeline from data collection to model development and analysis and online learning framework for modeling coil deflection at DIII-D. I will also briefly touch upon opportunities and challenges in development of digital twin framework.

Sammuli, Brian [General Atomics]

FAIRmaterials: Ontology Tools with Data FAIRification in Development

The bilingual FAIRmaterials package simplifies the creation and visualization of materials and data science ontologies. FAIRmaterials, available in the Python and R languages, addresses the complexities associated with traditional ontology editors based on manual user input such as Protege with an intuitive workflow and easy-to-use templates, making it accessible to users both experienced and inexperienced with ontologies. The FAIRmaterials package is its ability to programatically convert simple and structured CSV inputs into rich, well-defined ontologies. This capability is designed to support the findability, accessibility, interoperability, and reusability (FAIR) of research data and serve as a tool in the process of data FAIRification. Its additional features, such as automated ontology merging, static visualizations, and comprehensive documentation for outputs extend its utility, making it a valuable tool for any researcher engaged in knowledge management.

Bradley, Alexander Harding [Case Western Reserve U

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)