Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)

Open-Source Data for MAC-POSTS: Mobility Data Analytics Center - Prediction, Optimization, and Simulation Toolkit for Transportation Systems

MAC-POSTS (Mobility Data Analytics Center - Prediction, Optimization, and Simulation toolkit for Transportation Systems) is a toolkit for dynamic transportation network modeling. Developed by the Mobility Data Analytics Center (MAC) at Carnegie Mellon University, this package implements many classic dynamic transportation network models, as well as new models proposed by MAC members. It has served as one building block for many other models and research projects. As such, this package used to be treated as an internal research project of the MAC lab, and admittedly, the code base is messy, and the interface is hard to use. However, we are working hard to make it a generally usable and useful toolkit for dynamic transportation network modeling. We would really appreciate any feedback, comments, suggestions, or criticisms.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

A critical review on additive manufacturing of refractory alloys from a data analytics perspective- beyond nickel-based superalloys

Refractory alloys (RAs) are promising materials due to their exceptional physicochemical properties, but most research remains at the laboratory scale. For broader adoption, advancements in manufacturing are essential. Because their high stability makes conventional methods like machining and casting difficult, additive manufacturing (AM) is emerging as an effective approach for fabricating refractory alloy components. However, AM's repeated non-equilibrium thermal cycles introduce undesired features (e.g. defects, anisotropic microstructures, and residual stresses), which are magnified due to RAs’ unique properties. This paper comprehensively reviews the state-of-the-art methods of AM for refractory alloys. It explores data analytics techniques to establish design rules based on multi-fidelity experimental and computational methods. Furthermore, it investigates integrated, collaborative efforts to harmonise standalone databases, information, knowledge, and predictive models at multi-physics, multi-stage, and multi-scale. Unlike the existing literature that focuses primarily on material systems or process fundamentals, this work provides an integrated perspective on AM of refractory alloys from a data analytics standpoint, highlighting the roles of integrated computational materials engineering (ICME), verification, validation, and uncertainty quantification (VV&UQ), and digital twin-driven qualification in overcoming data scarcity and accelerating rapid qualification.

Additive manufacturing

INTEGRATION OF DATA ANALYTICS WITH SYSTEM HEALTH PROGRAMS

Industry equipment reliability and asset management programs are essential elements that help ensure the safe and economical operation of nuclear power plants. The effectiveness of these programs is addressed in several industry developed and regulatory programs. However, these programs have proven to be labor intensive and expensive. There is an opportunity to significantly enhance the collection, analysis, and use of this information to provide more cost-effective plant operation. Additionally, there is an acute industry need to leverage advanced technology to reduce costs and improve operational effectiveness. The goal of this paper is to provide effective and efficient analytical methods and tools to support risk-informed decisions for the equipment reliability and asset management programs at nuclear power plants. This is accomplished by creating a direct bridge between component health/lifecycle data and decision making (e.g., maintenance scheduling and project prioritization). Here we are supporting typical system engineer decisions regarding maintenance activity scheduling and component ageing management. This is performed in a risk-informed context where herein the term “risk” is broadly constructed to include both plant reliability and economics. This framework combines data analytics tools to analyze equipment reliability data with risk-informed methods designed to support system engineer decisions (e.g., maintenance and replacement schedules, optimal maintenance posture) in a customizable workflow. A challenge is that the structure of this workflow strongly depends on the decision that needs to be made, the type of data available, and the constraints that need to be considered. Current methods are designed to provide specific answers to specific problems; however, these methods might prove to be inadequate even when problem settings slightly change (e.g., different types of requirements, additional dependencies between system reliability and economics). We tackled this challenge by designing framework in a flexible and modular fashion such that the user can assemble and customize his/her own workflow that integrates SSC economic lifecycle models (e.g., maintenance and replacement costs), system reliability models, and optimization methods.

97 - MATHEMATICS AND COMPUTING

Monitoring L2 Milestone Summary: Converged HPC Center-Wide Data Analytics

This document is the summary for the ASC 2025 L2 milestone Converged HPC Center-Wide Data Analytics. It describes the convergence of the center-wide monitoring at LC, the data sources involved, and different ways this has enabled analyzing and gaining insights from the data.

97 MATHEMATICS AND COMPUTING

A Scientist-in-the-Loop Data Analytics Framework for Intelligent Simulation Model Tuning and Validation

This project developed a scientist-in-the-loop data analytics framework for intelligent simulation model tuning and validation, targeting the Weather Research and Forecasting (WRF) model and its solar energy variant, WRF-Solar-BNL. Domain experts, such as climate scientists, depend on large-scale numerical simulations for knowledge discovery and decision-making, yet the complexity of parameter tuning and the disconnect between automated optimization and domain expertise pose significant challenges. We extended an interactive visual analytics framework that enables domain experts to observe and intervene in the computational steering process by identifying disagreements between the simulation model, surrogate model, and the expert’s domain knowledge. Using Bayesian Optimization with Gaussian Process Regression as the surrogate model, our system allows users to probe parameter relationships, analyze correlation patterns, and adjust tuning parameters in real time. We developed use cases for solar irradiance forecasting through sustained collaboration with Brookhaven National Laboratory, resolving critical model configuration challenges and achieving meaningful reductions in prediction error. The project supported one PhD student, one MS student, and eight undergraduate students across three Data Science Capstone projects, resulting in one master’s thesis.

Dasgupta, Aritra [New Jersey Institute of Technolo

Predictive modeling of Néel temperature in austenitic alloys using CALPHAD and data analytics

The Néel temperature is a crucial yet often overlooked parameter in calculating the stacking fault energy (SFE) of austenitic alloys. Several empirical equations have been proposed to estimate the Néel temperature of austenitic alloys, which are then used to calculate the SFE and explain deformation mechanisms. However, these empirical equations, typically derived using linear regression algorithms, are often simplistic and may fail to capture the complex interactions among multiple alloying elements that influence the Néel temperature. Moreover, their applicability is usually limited to specific compositional ranges. In this study, we propose a CALPHAD based approach and develop a surrogate decision tree based regression model capable of capturing the interactions among multiple alloying elements to predict the Néel temperature. Predictions from both the CALPHAD approach and the regression model show close agreement with experimental measurements reported in the literature. In conclusion, the implications of accurate Néel temperature predictions on the calculated SFE and deformation mechanisms are also discussed.

36 MATERIALS SCIENCE

Machine Tool Data Analytics for Digital Twin and Machine Predictive Maintenance

The primary objective of this project is to improve machining process performance using in-process machining data from the machine tool controller and external sensors. Advances in the Industrial Internet of Things (IIoT) enable monitoring of machines using controller data. Examples of the data provided by a controller include execution status of the controller, part count, block of code being executed, door status, tool position, the spindle and axis load, etc. MTConnect and OPC-UA are the two common protocols for capturing machine information. In this collaboration, methods for retrieving the machine controller data from selected machine tool controls and making these data accessible in different subsystems (such as digital twins and machine maintenance portals, etc.) will be developed and tested. In addition, analytics to improve machining process performance (by increasing productivity and reducing downtime) will be developed.

42 ENGINEERING

Data Analytics Methods to Measure Plant Outage Resilience

Every 18 or 24 months nuclear power plants (depending on plant configuration, pressurized or boiling water reactor respectively) undergo a period of outage where the plant is taken offline and a large number of maintenance and surveillance activities (that cannot be performed while plant is running) are performed in typically 2–3 weeks. Planning of a plant outage is very challenging since all the activities are required to be performed in the shortest amount of time given available resources (typically contractor crews hired for the duration of the outage). Consequently, plant outages can be costly due the actual loss of power generation and crew costs and, because of it, there is a need to maximize resource usage in the outage planning phase and reduce the risk of outage delays. This paper is addressing these needs by providing a set of analytical methods designed to analyze plant outage schedule and identify critical elements based on available resources (time and crews). These methods are based on natural language processing and optimization algorithms. In this respect, two classes of methods have been developed: one that focuses on the time resource and how variability in the time to complete outage tasks may impact outage delays, and one that minimizes the risk of outage delays by integrating available resources to assess when daily activities should be performed.

97 - MATHEMATICS AND COMPUTING

Frameworks, Algorithms, and Scalable Technologies for Mathematics (FASTMath) SciDAC Institute

As computational models scale to larger computers, the rate at which they produce data has far outstripped the same computers ability to write that data and further the file systems ability to store that data. Almost all of the SciDAC applications, but especially those related to fusion solve very large scale PDEs whose scientific output his impacted by this problem. To gain access to dynamics in an exascale simulation that are not identifiable a priori and to make that dynamical data available to machine learning requires fundamental research in the area of in situ data data analytics. Here data analytics includes compression, visualization, uncertainty quantification, and machine learning. This in situ data analytics will enable on-the-fly spatial and temporal compression of solution dynamics, expose that space-time compressed field to machine learning algorithms that have been specialized to work with dynamically evolving data (existing machine learning algorithms treat data sets as static), greatly improving the opportunity for machine learning to provide feedback to the compression, all within an ongoing simulation, without the need to write data to files. The same concepts are also being applied to uncertainty quantification and multi-fidelity modeling which have similar needs for spatial and temporal compression of the ongoing exascale simulation to perform either without the typical, unacceptable writing of data to files.

97 MATHEMATICS AND COMPUTING

OEDI—Solar Grid Integration Data and Analytics Library

As a part of the Open Energy Data Initiative, this effort aims to develop and demonstrate novel distribution state estimation, control optimization, and transient analysis as well as provide access to data, data integration, and mapping information. More specifically, the focus of the effort will be on physics-based distribution system state estimation, hybrid (physics-based and machine learning) distribution optimal power flow, and event detection/analysis for solar integration and analytics. This work will enable reproducible, robust, replicable, and generalizable R&D in simulation and emulation of solar system integration. These test models and datasets will provide an integrated library for developing and testing power system operation technologies. To make the library user-friendly, this project will provide data curation tools such as data translators, mapping scripts and APIs, database schemas and metadata, interfaces and user dashboard, source code for the reference algorithms, description of the use-cases/scenarios, and comprehensive information on all the assumptions.

14 SOLAR ENERGY

A User-Facing Metric to Quantify the Quality of Mobility (CRADA Final Report)

The leading urban mobility data analytics firm StreetLight Data, Inc. partnered with the National Renewable Energy Laboratory to explore a commercial version of the Mobility Energy Productivity (MEP) metric. The commercialization effort was aimed to expand the adoption of the metric to key stakeholders in the urban planning space. Research comprised industry analysis, stakeholder feedback and conducting transportation practitioner focus groups. StreetLight concluded that commercialization of MEP is not feasible in the current market because users need a dynamic MEP tool that enables scenario planning. It should be able to calculate a MEP score dynamically (near instantaneous) when different inputs are changed.

33 ADVANCED PROPULSION SYSTEMS

Evaluation of Properties for Microsample Identification

A study was conducted to determine if individual particle characteristics could be used to identify particles of interest, sub-samples, from bulk post-detonation debris. Three archived post-detonation debris samples were used for this effort. Particles from these samples were identified as active (produced fission tracks), and inactive (did not produce fission tracks), as the first defining characteristic. Morphology was the secondary characteristic to select particles for further study, i.e. spherical/non-spherical. Once particles were identified and isolated, they were characterized by optical microscopy for size in µm, number of fission tracks, morphology, transmitted light color, and reflected light color. Particles were then analyzed by scanning electron microscopy for morphology, elemental content, and compound identification. Raman spectroscopy was attempted on five particles with indeterminate results due to environmental mixing (heterogeneity) during the events of particle formation. Once all non-destructive analyses were completed all particles were analyzed by thermal ionization mass spectrometry to determine isotopic atom percents of plutonium and uranium, and an estimate of atoms of plutonium and uranium in each particle. An estimate of the ratio of uranium to plutonium was also obtained (U/Pu). Data analytics of the data from the particles showed that combining characteristics of the particles have a high probability of identifying particles of interest from bulk post-detonation debris samples. Please note that this version of the report is an abridged version of the full report (Wagnon et al. 2025) that has been edited to be appropriate for public release.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O