Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Modern data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Advanced data science toolkit for non-data scientists – A user guide

Emerging modern data analytics attracts much attention in materials research and shows great potential for enabling data-driven design. Data populated from the high-throughput CALPHAD approach enables researchers to better understand underlying mechanisms and to facilitate novel hypotheses generation, but the increasing volume of data makes the analysis extremely challenging. Here in this paper, we introduce an easy-to-use, versatile, and open-source data analytics frontend, ASCENDS (Advanced data SCiENce toolkit for Non-Data Scientists), designed with the intent of accelerating data-driven materials research and development. The toolkit is also of value beyond materials science as it can analyze the correlation between input features and target values, train machine learning models, and make predictions from the trained surrogate models of any scientific dataset. Various algorithms implemented in ASCENDS allow users performing quantified correlation analyses and supervised machine learning to explore any datasets of interest without extensive computing and data science background. The detailed usage of ASCENDS is introduced with an example of experimental high-temperature alloy data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

Artificial Intelligence for Earth System Predictability (AI4ESP) (2021 Workshop Report)

In October 2021, the U.S. Department of Energy (DOE) welcomed participants to the Artificial Intelligence for Earth System Predictability (AI4ESP) Workshop, hosted by the Office of Biological and Environmental Research (BER)—Advanced Scientific Computing Research (ASCR). The workshop is part of BER-ASCR’s ambition to more radically and aggressively advance prediction capabilities in the climate, Earth, and environmental sciences through the use of modern data analytics and artificial intelligence (AI). Advances in these capabilities are needed to improve predictions of climate change and extreme events that provide actionable information for planning and building resilience to their impacts.

54 ENVIRONMENTAL SCIENCES↗

Emulation and detection of physical faults and cyber-attacks on building energy systems through real-time hardware-in-the-loop experiments

The increasing use of remote or mobile access, integrated wearable technologies, data exchange, and cloud-based data analytics in modern smart buildings is steering the building industry towards open communication technologies. The increased connectivity and accessibility could lead to more cyber-attacks in smart buildings. On the other hand, physical faults (e.g., HVAC -heating, ventilation, and air-conditioning faults) may have similar adverse impacts as those from the cyber-attacks on building energy systems, such as occupant discomfort, energy wastage, and equipment downtime. However, current physical behavior-based anomaly detection methods fail to differentiate between cyber-attacks and physical faults in building energy systems. Moreover, the challenge in collecting real-world threat data with ground truth has led researchers to rely on numerical models with user-defined assumptions, which may not accurately reflect real-world conditions due to the lack of in-situ experimental datasets. To address these challenges and gaps, this paper presents a flexible hardware-in-the-loop (HIL) testbed for generating cyber-attack and physical fault datasets and demonstrating threat detection algorithms in a real building automation system (BAS) environment. This testbed combines hardware (i.e., real BAS with local HVAC controllers and a physical network) with software (i.e., high-fidelity models to represent behaviors of building envelope and HVAC energy systems), enabling emulations of realistic threats. Five HIL experiments, including one baseline without any threats, two with physical faults, and two with cyber-attacks, were conducted to generate datasets containing detailed network traffic and system states. A joint classification framework, incorporating a network analyzer and a physical HVAC fault detector, was proposed to automatically detect cyber-physical abnormalities on BAS at both the network and the physical HVAC levels. The network analyzer comprises a conditional random fields (CRF) based command validator and a statistics-based detection strategy. The fault detector employs a weather and schedule-based pattern matching and feature-based principal component analysis (WPM-FPCA) method. Evaluation of the classification using four metrics from the multi-class confusion matrix revealed an average accuracy of 90.2%, recall of 89.7%, precision of 88.5% and F1-score of 89.2%. Finally, these results demonstrate that the proposed joint classification framework can effectively differentiate between specific types of cyber-attacks (e.g., device reinitialization attack, network Denial-of-Service attack) and physical faults (e.g., air handling unit operational fault, cooling coil valve stuck) in real time for improved building energy management.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

Real-time Data Analytics for Condition Monitoring of Complex Industrial Systems

Modern industrial systems are now fitted with several sensors for condition monitoring. This is advantageous because these sensors can provide mass amounts of data that have the potential for aiding in tasks such as fault detection, diagnosis, and prognostics. However, the information valuable for performing these tasks is often clouded in noise and must be mined from high-dimensional data structures. Therefore, this dissertation presents a data analytics framework for performing these condition monitoring tasks using high-dimensional data. Demonstrations of this framework are detailed for challenges related to power generation systems in automobiles, power plants, and aircraft engines. These implementations leverage data collected from state-of-the-art, industry class test-rigs. Results indicate the ability of this framework to develop effective methodologies for condition monitoring of complex systems.

Peters, Benjamin↗

FY22 Laboratory Directed Research and Development Annual Report

The Laboratory Directed Research and Development (LDRD) program yields foundational scientific research and development (R&D) essential to growing SRNL’s core competencies, in alignment with SRNL’s Strategic Plan to provide long-term benefits to the Department of Energy (DOE), the National Nuclear Security Administration (NNSA), and other customers and stakeholders. Five strategic goals are outlined in SRNL’s strategic plan: 1) Provide applied science and engineering for EM’s active clean-up sites and LM’s post closure management sites; 2) Provide science-based solutions for gaps identified in nonproliferation strategic vision and support the government in actives impacting national security; 3) Lead ST&E as the central technical authority for processing tritium loaded reservoirs and support production of plutonium pits; 4) Align science and energy security programs by focusing modern modeling, simulation, and data analytics tools on materials engineering and performance applications; 5) Build a workforce for the future.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Singleton Sieving: Overcoming the Memory/Speed Trade-Off in Exascale k-mer Analysis

Traditional filter data structures, such as Bloom filters, do not offer necessary features that modern high-performance data analytics applications need in order to efficiently perform complex data analysis tasks. For example, MetaHipMer, a de novo metagenome assembler, can use filters to weed out singleton k-mers and reduce memory usage by 30%-70%. However, the filter needs the ability to associate values with k-mers in order to perform the analysis in a single communication pass. Bloom filters do not support value associations and cause the application to perform an extra communication pass, thereby increasing the run time. Therefore, MetaHipMer faces a trade off between memory and speed due to the limited capabilities of traditional filters. In this paper, we overcome the memory and speed trade off in MetaHipMer by integrating a GPU-based feature-rich filter, the Two-Choice filter (TCF), in the MetaHipMer pipeline. The TCF uses key-value association to approximately store k-mers with extensions. This allows MetaHipMer to perform k-mer analysis on the GPUs in a single communication pass. Our empirical analysis shows a 50% reduction in memory usage in k-mer analysis on each node in MetaHipMer without any effect on the overall run time or assembly quality. The memory reduction in turn results in a 43% reduction in the number of nodes required to assemble datasets and enables MetaHipMer to scale to much larger datasets.

McCoy, Hunter↗

FY24 Laboratory Directed Research and Development Annual Report

The Laboratory Directed Research and Development (LDRD) program yields foundational scientific research and development (R&D) essential to growing SRNL’s core competencies, in alignment with SRNL’s Strategic Plan to provide long-term benefits to the Department of Energy (DOE), the National Nuclear Security Administration (NNSA), and other customers and stakeholders. Five strategic goals are outlined in SRNL’s strategic plan: 1) Provide applied science and engineering for EM’s active clean-up sites and LM’s post closure management sites 2) Provide science-based solutions for gaps identified in nonproliferation strategic vision and support the government in activities impacting national security 3) Lead Science, Technology & Engineering as the central technical authority for processing tritium loaded reservoirs and support production of plutonium pits 4) Align science and energy security programs by focusing modern modeling, simulation, and data analytics tools on materials engineering and performance applications 5) Build a workforce for the future

Clark, Sue [Savannah River National Laboratory (SR↗

Leveraging Cloud Platforms for Grid Modernization

Presentation held on Friday December 5th, 2025 at the “San Diego Tech Conference and Expo” about “Advanced Sensor Data Analytics and Cloud Computation for Grid Modernization”

24 POWER TRANSMISSION AND DISTRIBUTION↗

Structure-aware graph neural network based deep transfer learning framework for enhanced predictive analytics on diverse materials datasets

Abstract Modern data mining methods have demonstrated effectiveness in comprehending and predicting materials properties. An essential component in the process of materials discovery is to know which material(s) will possess desirable properties. For many materials properties, performing experiments and density functional theory computations are costly and time-consuming. Hence, it is challenging to build accurate predictive models for such properties using conventional data mining methods due to the small amount of available data. Here we present a framework for materials property prediction tasks using structure information that leverages graph neural network-based architecture along with deep-transfer-learning techniques to drastically improve the model’s predictive ability on diverse materials (3D/2D, inorganic/organic, computational/experimental) data. We evaluated the proposed framework in cross-property and cross-materials class scenarios using 115 datasets to find that transfer learning models outperform the models trained from scratch in 104 cases, i.e., ≈90%, with additional benefits in performance for extrapolation problems. We believe the proposed framework can be widely useful in accelerating materials discovery in materials science.

Chemistry↗

First measurement of large area jet transverse momentum spectra in heavy-ion collisions

Jet production in lead-lead (PbPb) and proton-proton (pp) collisions at a nucleon-nucleon center-of-mass energy of 5.02 TeV is studied with the CMS detector at the LHC, using PbPb and pp data samples corresponding to integrated luminosities of 404 μb$^{−1}$ and 27.4 pb$^{−1}$, respectively. Jets with different areas are reconstructed using the anti-k$_{T}$ algorithm by varying the distance parameter R. The measurements are performed using jets with transverse momenta (p$_{T}$) greater than 200 GeV and in a pseudorapidity range of |η| < 2. To reveal the medium modification of the jet spectra in PbPb collisions, the properly normalized ratio of spectra from PbPb and pp data is used to extract jet nuclear modification factors as functions of the PbPb collision centrality, p$_{T}$ and, for the first time, as a function of R up to 1.0. For the most central collisions, a strong suppression is observed for high-p$_{T}$ jets reconstructed with all distance parameters, implying that a significant amount of jet energy is scattered to large angles. The dependence of jet suppression on R is expected to be sensitive to both the jet energy loss mechanism and the medium response, and so the data are compared to several modern event generators and analytic calculations. The models considered do not fully reproduce the data.[graphic not available: see fulltext]

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Oak Ridge National Laboratory NCSP Analytical Methods Subtask 3, AMPX Development and Maintenance, and NCSP Nuclear Data Subtask 6, SAMMY Modernization

The modernization of SAMMY continued with consolidation of access to covariance information for adjusted parameters and data in SAMMY. This consolidation allowed for removal of many scratch files. In addition, work was initiated to make the 0K cross section calculation more modular and less dependent on SAMMY global parameters. An initial application programming interface (API) was added to expose cross sections (including resolution broadening) generated by SAMMY to external fitting routines. The processing for thermal moderators in AMPX was updated for selected moderators for which the generated grid was not fine enough. Updated libraries were generated for SCALE. In addition, work continued to fully support new Evaluated Nuclear Data File (ENDF) formats, including the Generalized Nuclear Database Structure (GNDS) in AMPX.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

NCSP Analytical Methods Subtask 3, AMPX Development and Maintenance, and NCSP Nuclear Data Subtask 6, SAMMY Modernization

The modernization of SAMMY continued with the removal of all container array access for the main program. The container array was used due to the computer limitations at the time of the initial writing of SAMMY, but it prevented the use of some modern code testing options. Additionally, the authors continued to make the code more modular by shifting the storage of the energy grid and the covariance information to C++ storage, thus allowing for better unit testing and the elimination of scratch files. For AMPX, the authors continued to implement the new Evaluated Nuclear Data File format to allow AMPX to read new evaluations as they become available.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Guest Editorial: Advanced Data-Analytics for Power System Operation, Control, and Enhanced Situational Awareness

Along with the smart grid development, modern power systems are entering a ‘data-intensive’ era. A vast volume of data from power grids is being collected through advanced sensing and communication technologies, such as smart metering data, phasor measurement data, as well as meteorological data (e.g., wind speed and solar irradiance) related to renewable power generation. Such data contains comprehensive information about the power system covering equipment's health status, power grid's static and dynamic characteristics, renewable power generation, customers’ electricity usage pattern, etc. Therefore, advanced data-analytics techniques are needed to convert such data to knowledge for practical applications. In line with the trend of widespread data-driven applications in power systems, this Special Issue aims to present state-of-the-art research works on advanced data-analytics for power system's operation, control, and situational awareness. There are in total twenty-six papers accepted for publication in this Special Issue through careful peer reviews and revisions. Under the overarching theme of data-driven applications in power systems, the selected papers are broadly categorised into five topics. The summary of every topic is given below. You are, however, strongly encouraged to read the full paper if interested.

Xu, Yan↗

Data Analytics and Visualization of Energy Systems for Critical Infrastructure Insights

Modernization of energy systems including transportation facilities provides opportunities for increased efficiency, expansion of commerce and meeting industry and federal goals. A significant increase in electrical demand is projected to meet these needs, which concentrates at facilities such as airports. For example, Xcel Energy working with two airports in their service area recently published information projecting an up to fivefold increase in electricity demand in the next 25 years [1]. Concurrently, the US Government Accountability Office (GAO) recently surveyed 30 commercial service airports identifying more than 300 outages of more than 5 minutes between 2015 and 2022 [2]. Power, reliability, and resilience planning becomes more important to safely maintain operations and the flow of commerce with fewer energy carriers providing necessary energy to safely move passengers and goods. NREL proposes to develop methodologies to allow owners, utilities, and federal agencies to dynamically analyze, forecast, and manage energy loads at airports, focused upon maintaining the flow of commerce in an efficient, sustainable, and resilient way. To address these energy challenges, a suite of technologies and methodologies can be leveraged to validate concepts, inform design, de-risk solutions and optimize energy management during deployment. These technologies include digitalization of energy systems, microgrid methodologies, and related energy technologies for building and vehicle loads. [1] Electrifying Airport Ecosystems - https://www.enterprisemobility.com/content/dam/enterpriseholdings/marketing/innovation-in-mobility/vehicle-innovation/airport-electrification-study-full-report-2024.pdf [2] Airport Infrastructure: Selected Airport's Efforts to Enhance Electrical Resilience https://www.gao.gov/products/gao-23-105203.

critcal infrastructure↗

Large-Scale Trajectory Analysis via Feature Vectors

The explosion of both sensors and GPS-enabled devices has resulted in position/time data being the next big frontier for data analytics. However, many of the problems associated with large numbers of trajectories do not necessarily have an analog with many of the historic big-data applications such as text and image analysis. Modern trajectory analytics exploits much of the cutting-edge research in machine-learning, statistics, computational geometry and other disciplines. We will show that for doing trajectory analytics at scale, it is necessary to fundamentally change the way the information is represented through a feature-vector approach. We then demonstrate the ability to solve large trajectory analytics problems using this representation.

58 GEOSCIENCES↗