Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Plutonium isotope ratio measurements by total evaporation-thermal ionization mass spectrometry (TE-TIMS): an evaluation of uncertainties using traceable standards from the New Brunswick Laboratory

The accuracy and precision of isotope amount ratio measurements using thermal ionization mass spectrometry (TIMS) instrumentation are described, and the measurement of Pu materials is emphasized. The mass fractionation observed for Am, Ga, Pu, and U for isotope amount ratio measurements using the total evaporation (TE) technique is compared with theoretical estimates to demonstrate the advantage of the TE methodology and to investigate systematic biases in the major isotope amount ratios of U and Pu certified reference material (CRM) standards from the U.S. provider of CRMs. The quality of the Pu isotopic data generated by TIMS instruments in an analytical laboratory is demonstrated by the application of the double ratio technique to estimate the 241 Pu half-life. Analytical data on traceable Pu CRMs from the New Brunswick Laboratory (NBL), generated as part of routine measurements supporting various programs, are used for this half-life estimation. Although the 241 Pu abundances in CRMs of 136, 137, 138, and 126-A are approximately 200–2000× smaller than those in the 241 Pu material used in the previous Institute for Reference Materials and Measurements (IRMM) evaluation of the 241 Pu half-life, the half-life value estimated in this work shows excellent agreement with the currently accepted value from the IRMM. This agreement also demonstrates the pedigree of the Pu isotopic standards from the NBL and the quality of the isotope amount ratio measurements using TIMS instrumentation. For both the major and minor Pu isotope amount ratios, this report describes the relative importance of the factors affecting the uncertainty of TIMS measurements, which are considered the gold standard in isotope ratio measurements (LA-UR-24-29199).

07 ISOTOPE AND RADIATION SOURCES

Physics-coupled data-driven design of high-temperature alloys

We present a materials design loop, which streamlines physics-coupled machine learning (ML) surrogate models to discover new alloy chemistries with improved properties. The efficacy is demonstrated by discovering a high-temperature alumina-forming austenitic (AFA) stainless steel with enhanced creep, followed by experimental validation. The ML models have been trained using a well-curated, highly consistent experimental dataset augmented with synthetic microstructural features from a computational thermodynamic approach. We have populated a large number of hypothetical AFA alloys to explore the high-dimensional composition space and have predicted their creep properties by providing the same synthetic input features obtained from the trained ML models. Uncertainties from the ML training were taken as thresholds for truncating predicted results to identify alloys with improved or deteriorated creep. Individual elemental compositions have been determined via probability density distribution analysis from the group of alloys at the top and bottom of the predicted creep values for further virtual and experimental validations. In conclusion, we anticipate that this workflow can be applied to screen desired conditions, such as chemistry and processing parameters, in high-dimensional space through physics-guided data analytics.

Alloy design

Quantifying Uncertainty in HPC Job Queue Time Predictions

High Performance Computing (HPC) has developed at an unprecedented pace in recent decades. This growth has demanded corresponding development in the area of HPC Operational Data Analytics (ODA), which encompasses a wide range of data analysis techniques, ML/AI efforts, tools, and visualizations. Published studies in ODA offer a variety of practical ways to inform HPC users, administrators, procurement managers, and other stakeholders. Uncertainty analysis, however, is rare in the related published literature. For instance, we identify only 1 out of 14 existing studies focused on job queue time prediction that investigates the uncertainty aspect of their proposed predictions. We recognize the utmost importance uncertainty quantification can have in such predictive analytics solutions, with consequences in how users interpret information they receive, and attempt to bridge this gap. With the goal of improving access to such insights, we develop a process for determining upper and lower bounds of the predicted queue times of a regression model at a specified confidence level. Our current research is focused on the uncertainty in predicting job queue times, yet our approach may be employed in predicting other metrics.

HPC

The role of quantum computing in advancing scientific high-performance computing: A perspective from the ADAC institute

Quantum computing (QC) has gained significant attention over the past two decades due to its potential for speeding up classically demanding tasks. This transition from an academic focus to a thriving commercial sector is reflected in substantial global investments. While advancements in qubit counts and functionalities continue at a rapid pace, current quantum systems still lack the scalability for practical applications, facing challenges such as too high error rates and limited coherence times. Here, this perspective paper examines the relationship between QC and high-performance computing (HPC), highlighting their complementary roles in enhancing computational efficiency. It is widely acknowledged that even fully error-corrected QC will not be suited for all computational tasks. Rather, future compute infrastructures are anticipated to employ quantum acceleration within hybrid systems that integrate HPC and QC. While QC can enhance classical computing, traditional HPC remains essential for maximizing quantum acceleration. This integration is a priority for supercomputing centers and companies, sparking innovation to address the challenges of merging these technologies. The novelty of this work lies in its unique perspective, reflecting the collective insights of the Accelerated Data Analytics and Computing (ADAC) Institute, a global consortium of over 20 leading HPC centers. Recognizing the growing importance of QC, ADAC established a Quantum Computing Working Group in 2023 to foster collaboration and knowledge-sharing among its members. This paper synthesizes insights from the group’s collaborative efforts and incorporates findings from a member survey that captures shared experiences, ongoing projects, and strategic directions. By outlining the current landscape and challenges of QC integration into HPC ecosystems, this work offers HPC specialists practical and forward-looking guidance on the opportunities and implications of QC in computationally intensive endeavors.

Accelerated Data Analytics and

A path to intelligent watersheds: coordinating the data to decision pipeline

Operations of multi-reservoir systems are challenged in-part by the interplay of complex physical processes functioning within the watershed. The employment of intelligent systems can be of aid by linking environmental sensing, information technology, data analytics, simulation and decision support to achieve a data-to-decision flow of information. A further challenge is that watershed resources are managed for multiple purposes requiring some level of coordination among numerous resource managers, asset operators and users. System intelligence in this context relies on shared community platforms (data portals, community models), and coordinated communication between decision makers. Opportunities to enrich watershed intelligence has been the subject of a roadmapping exercise for the Department of Energy’s Water Power Technologies Office which has relied on broad stakeholder engagement. Initial phases of engagement involved personal interviews and a series of virtual group meetings, which focused on identifying opportunities to improve the intelligence of the physical infrastructure within our watersheds—examples of feedback include improved sensing of snowpack and runoff, data standards for facilitated data sharing, and better forecasting tools. The latter phase of engagement involved the conduct of a case study in the Upper Colorado River basin where key stakeholders were interviewed to map how their decisions are informed by intelligence from other basin stakeholders. Our presentation will highlight the interdisciplinary flow of information in complex watershed systems and identify physical and institutional opportunities toward the strategic operation of water infrastructure.

Colorado River

EDX ClaiMM

EDX ClaiMM is a centralized data & analytical platform designed to revolutionize U.S. critical minerals and materials (CMM) activities. By providing a robust digital infrastructure, ClaiMM will accelerate the combination, leveraging, and rapid utilization of vital data, advanced tools, and cutting-edge research advancements in CMM. This adaptive digital research hub connects the CMM community to essential knowledge products and offers access to interoperable datasets, databases, models, software, and tools from the National Energy Technology’s (NETL’s) Energy Data eXchange (EDX) and other authoritative sources, serving both public and private sectors. EDX ClaiMM delivers AI-informed solutions to address fundamental knowledge gaps and fosters the innovation of new techniques for enhanced characterization and recovery of CMMs within the U.S. By leveraging cloud-hosted, scalable digital infrastructure, ClaiMM meets public–private applied energy needs. It equips the CMM community with priority digital resources that harness on-site and cloud compute capabilities, enabling big data storage, advanced processing, analytics, and visualization.

Critical Materials; Critical Minerals; Rare Earth

Hazard Detection Detector Cards

This report presents a comprehensive summary of five advanced anomaly detection tools developed and deployed by Oak Ridge National Laboratory in support of the VA’s Health Information Technology modernization. These detectors—Order Path Tracker, Trend Watcher, Pain Pointer, Performance Monitor, and Patient Record Flag Detector—leverage statistical and machine learning methods to monitor workflow disruptions, detect anomalies in care sequences and volumes, identify bottlenecks, and track system-level performance metrics across VistA and Millennium systems. All detectors have been integrated into the Health Data Analytics Platform (HDAP), with most having completed deployment and testing using live data from targeted stations in cardiology and oncology domains. This work enhances VA’s capacity for proactive system surveillance, promotes patient safety, and informs data-driven operational improvements across the EHR ecosystem.

97 MATHEMATICS AND COMPUTING

Data Quality Monitoring for the Hadron Calorimeters Using Transfer Learning for Anomaly Detection

The proliferation of sensors brings an immense volume of spatio-temporal (ST) data in many domains, including monitoring, diagnostics, and prognostics applications. Data curation is a time-consuming process for a large volume of data, making it challenging and expensive to deploy data analytics platforms in new environments. Transfer learning (TL) mechanisms promise to mitigate data sparsity and model complexity by utilizing pre-trained models for a new task. Despite the triumph of TL in fields like computer vision and natural language processing, efforts on complex ST models for anomaly detection (AD) applications are limited. In this study, we present the potential of TL within the context of high-dimensional ST AD with a hybrid autoencoder architecture, incorporating convolutional, graph, and recurrent neural networks. Motivated by the need for improved model accuracy and robustness, particularly in scenarios with limited training data on systems with thousands of sensors, this research investigates the transferability of models trained on different sections of the Hadron Calorimeter of the Compact Muon Solenoid experiment at CERN. The key contributions of the study include exploring TL’s potential and limitations within the context of encoder and decoder networks, revealing insights into model initialization and training configurations that enhance performance while substantially reducing trainable parameters and mitigating data contamination effects.

47 OTHER INSTRUMENTATION

Reducing Data Center Peak Cooling Demand and Energy Costs with Underground Thermal Energy Storage (UTES)

By recent estimates, data center energy demands are projected to consume between 6.7% and 12% of U.S. annual electricity generation by the year 2028, driven primarily by expanded demands from cloud services, big data analytics, and Artificial Intelligence (AI) (Shehabi et al., 2024). As much as 40% of data center total energy consumption are loads associated with the site infrastructure cooling systems, and these are often highly water consumptive (Aljbour et al., 2024). For energy system planners, this presents significant challenges to meeting and managing the anticipated loads, and especially the peak loads of projected data center deployments. Geothermal technologies offer two unique solutions to these challenges: 1) by serving loads through the deployment of new conventional and/or next-generation geothermal power technologies such as EGS and 2) through an often-overlooked opportunity to reduce data center peak cooling loads. The latter is the focus of this paper which explores Cold Underground Thermal Energy Storage ("Cold UTES") as an emerging industrial-scale geothermal cooling solution. This cooling solution is energy efficient, non-water-consumptive, and utilizes long duration energy storage (LDES) on both diurnal and seasonal time scales. Cold UTES has the potential to also function as a virtual power plant (VPP). The US Department of Energy's Geothermal Technologies Office is supporting R&D to understand the grid and system-wide value, costs, and impacts of deploying this emergent cooling solution at scale.

AI

A Co-Registered In-Situ and Ex-Situ Dataset of Electrical, Acoustic, and CT Characteristics from Wire Arc Additive Manufacturing Process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data centric approach emphasizes leveraging available data throughout the production process to optimize performance. Integration of extensive data analysis provides the opportunity to improve precision, reduce waste, and enhance the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes the comprehensive description of deposition process, process parameters, in-situ collected welding characteristics, acoustic data, and X-Ray Computed Tomography analysis data for the build. Dataset A Co-Registered In-Situ and Ex-Situ Dataset of Electrical, Acoustic, and CT Characteristics from Wire Arc Additive Manufacturing Process has arisen under UT-Battelle, LLC’s Prime Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy (DOE) to manage and operate the Oak Ridge National Laboratory. UT-Battelle, LLC will not assert any rights under United States law or under the Prime Contract it has in the dataset against any user of the dataset, including any copyrights or patent rights. UT-Battelle, LLC requests that attribution to the dataset is provided as academically appropriate.

42 ENGINEERING

Data for Machado-Silva et al. (2024), "Short-Term Groundwater Level Fluctuations Drive Subsurface Redox Variability"

This dataset contains the analytical data reported in Machado-Silva et al. (2024) as part of the COMPASS-FME project, which seeks to advance a scalable, predictive understanding of the fundamental biogeochemical processes, ecological structure, and ecosystem dynamics that distinguish coastal terrestrial-aquatic interfaces from the purely terrestrial or aquatic systems to which they are coupled. The dataset consists of water quality parameters as well as redox potential, water content, and electrical conductivity. These data were collected in 2022 in Crane Creek (CRC), Portage River (PTR), and Old Woman Creek (OWC). Each of these sites included uplands (UP), transitions (TR), wetland-transition edge (WTE), and wetland (W) zones. The sites represent replicates of the Lake Erie terrestrial-aquatic interface under fluctuating water levels and are located in well-preserved areas with natural or restored marsh and forest cover.This dataset consists of a single data file (Machado_Silva_et_al_2024_EST_data.csv) that is in comma-separated value (CSV) format. No special software is required to read it.This dataset uses the ESS-DIVE Hydrologic Monitoring Reporting Format 1.0.

54 ENVIRONMENTAL SCIENCES

A co-registered in-situ and ex-situ dataset from wire arc additive manufacturing process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data-centric approach emphasizes leveraging sensor data available throughout the production process to optimize performance. Integration of extensive data analysis provides opportunities for improving precision, reducing waste, and enhancing the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes a comprehensive description of the deposition process, process parameters, welding characteristics and acoustic data collected in-situ, and X-Ray Computed Tomography data of the build.

42 ENGINEERING

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING

Roadmap on data-centric materials science

Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

36 MATERIALS SCIENCE

Hyperspectral segmentation of plants in fabricated ecosystems

Hyperspectral imaging provides a powerful tool for analyzing above-ground plant characteristics in fabricated ecosystems, offering rich spectral information across diverse wavelengths. This study presents an efficient workflow for hyperspectral data segmentation and subsequent data analytics, minimizing the need for user annotation through the use of ensembles of sparse mixed scale convolution neural networks. The segmentation process leverages the diversity of ensembles to achieve high accuracy with minimal labeled data, reducing labor-intensive annotation efforts. To further enhance robustness, we incorporate image alignment techniques to address spatial variability in the dataset. Downstream analysis focuses on using the segmented data for processing spectral data, enabling monitoring of plant health. This approach provides a scalable solution for spectral segmentation, and facilitates actionable insights into plant conditions in complex, controlled environments. Our results demonstrate the utility of combining advanced machine learning techniques with hyperspectral analytics for high-throughput plant monitoring.

Zwart, Petrus H.

A Performance Model of In-Situ Techniques

The computational capacity of High-Performance Computing (HPC) systems increases continuously with the rapid development of central processing units (CPUs) and graphic processing units (GPUs), while the in-/output (IO) subsystem develops relatively slowly and storage capacity is also limited. Data-intensive applications, which are designed to leverage the high computational capacity of HPC resources, typically generate a considerable amount of data for post-processing visualizations and data analytics. The limited IO speed and storage space could lead to constraints in the actual performance of these applications and, therefore, scientific discovery. In-situ techniques, where data is visualized/analysed while still in memory rather than through disk, can contribute to alleviating these problems as they can reduce or even fully avoid data writing/reading through the IO subsystem to/from storage. However, the overall efficiency of insitu techniques crucially depends on the characteristics of both the in-situ tasks and the applications, and the resource distribution among them. Therefore, choosing the right in-situ approach (synchronous, asynchronous, or hybrid) and resource allocation is essential to minimize overhead and maximize the benefits of concurrent execution. In this paper, we present a performance model of in-situ techniques to find the most beneficial in-situ approach and the preferred resource configuration. We verify the high accuracy of our approach with over 6800 measurements and provide use cases with different applications.

Ju, Yi [Max Planck Computing and Data Facility, Ga

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics