Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Training Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

End-to-end online performance data capture and analysis for scientific workflows

With the increased prevalence of employing workflows for scientific computing and a push towards exascale computing, it has become paramount that we are able to analyze characteristics of scientific applications to better understand their impact on the underlying infrastructure and vice-versa. Such analysis can help drive the design, development, and optimization of these next generation systems and solutions. Here, we present the architecture, integrated with existing well-established and newly developed tools, to collect online performance statistics of workflow executions from various, heterogeneous sources and publish them in a distributed database (Elasticsearch). Using this architecture, we are able to correlate online workflow performance data, with data from the underlying infrastructure, and present them in a useful and intuitive way via an online dashboard. We have validated our approach by executing two classes of real-world workflows, both under normal and anomalous conditions. The first is an I/O-intensive genome analysis workflow; the second, a CPU- and memory-intensive material science workflow. Based on the data collected in Elasticsearch, we are able to demonstrate that we can correctly identify anomalies that we injected. The resulting end-to-end data collection of workflow performance data is an important resource of training data for automated machine learning analysis.

97 MATHEMATICS AND COMPUTING↗

A Novel Machine Learning Approach to Disentangle Multitemperature Regions in Galaxy Clusters

The hot intracluster medium (ICM) surrounding the heart of galaxy clusters is a complex medium that comprises various emitting components. Although previous studies of nearby galaxy clusters, such as the Perseus, the Coma, or the Virgo cluster, have demonstrated the need for multiple thermal components when spectroscopically fitting the ICM’s X-ray emission, no systematic methodology for calculating the number of underlying components currently exists. In turn, underestimating or overestimating the number of components can cause systematic errors in the emission parameter estimations. In this paper, we present a novel approach to determining the number of components using an amalgam of machine learning techniques. Synthetic spectra containing a various number of underlying thermal components were created using well-established tools available from the Chandra X-ray Observatory. The dimensions of the training set was initially reduced using principal component analysis and then categorized based on the number of underlying components using a random forest classifier. Our trained and tested algorithm was subsequently applied to Chandra X-ray observations of the Perseus cluster. Our results demonstrate that machine learning techniques can efficiently and reliably estimate the number of underlying thermal components in the spectra of galaxy clusters, regardless of the thermal model (MEKAL versus APEC). We also confirm that the core of the Perseus cluster contains a mix of differing underlying thermal components. We emphasize that although this methodology was trained and applied on Chandra X-ray observations, it is readily portable to other current (e.g., XMM-Newton, eROSITA) and upcoming (e.g., Athena, Lynx, XRISM) X-ray telescopes. The code is publicly available at https://github.com/XtraAstronomy/Pumpkin.

79 ASTRONOMY AND ASTROPHYSICS↗

Preliminary report on applications of machine learning techniques to the Nevada play fairway analysis

We are applying machine learning (ML) techniques, including training set augmentation and artificial neural networks, to mitigate key challenges in the Nevada play fairway project. The study area includes ~85 active geothermal systems as potential training sites and >12 geologic, geophysical, and geochemical features. The main goal is to develop an algorithmic approach to identify new geothermal systems in the Great Basin region. Major objectives include: 1) integrate ML techniques into the geothermal community; 2) develop open community datasets, whereby all play fairway and ML datasets and algorithms are publicly released and available for modification by various user groups; 3) identify data acquisition targets with high value for future work; 4) identify new signatures to detect blind geothermal systems; and 5) foster new capabilities for characterizing subsurface temperature and permeability. Initially, ML techniques are being applied to the same play fairway datasets and workflow. ML will then be applied to both enhanced and additional datasets, with modification of the PFA workflow to incorporate the new datasets. Finally, ML will be applied to define new workflows using the enhanced and additional datasets. An algorithmic approach that empirically learns to estimate weights of influence for diverse parameters can potentially scale and perform better than the play fairway analysis. Initial work on this project has involved 1) evaluating potential positive and negative training sites, 2) transformation of datasets into formats suitable for ML, and 3) initial development and testing of ML techniques.

58 GEOSCIENCES↗

Training toward significance with the decorrelated event classifier transformer neural network

Experimental particle physics uses machine learning for many tasks, where one application is to classify signal and background events. This classification can be used to bin an analysis region to enhance the expected significance for a mass resonance search. In natural language processing, one of the leading neural network architectures is the transformer. In this work, an event classifier transformer is proposed to bin an analysis region, in which the network is trained with special techniques. The techniques developed here can enhance the significance and reduce the correlation between the network’s output and the reconstructed mass. It is found that this trained network can perform better than boosted decision trees and feed-forward networks. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Clean Energy Education and Training Resources and Opportunities in New York's Southern Tier Region

New York's Southern Tier Region is experiencing high growth and investment in the clean energy sector and is anticipating more jobs to come in energy efficiency, renewable energy, and manufacturing in the coming years. The Network for a Sustainable Tomorrow (NEST) is a nonprofit network of programs working to develop a regional backbone system for education and training programs as well as curricula to support the workforce needed for the region's growing industries to succeed. Through its participation in the US Department of Energy's Better Buildings Workforce Accelerator, NEST requested technical assistance in conducting a landscape and needs assessment of the region's existing clean energy education and workforce development assets. This report supports NEST's efforts by providing a baseline of clean energy employment data, an inventory and gap analysis of the education and workforce development assets currently available and serving the Southern Tier Region, and case studies of innovative and successful regional clean energy education and workforce coalitions from around the county.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Towards Redefining the Reproducibility in Quantum Computing: A Data Analysis Approach on NISQ Devices

Although the building of quantum computers has kept making rapid progress in recent years, noise is still the main challenge for any application to leverage the power of quantum computing. Existing works addressing noise in quantum devices proposed noise reduction when deploying a quantum algorithm to a specified quantum computer. The reproducibility issue of quantum algorithms has been raised since the noise levels vary on different quantum computers. Importantly, existing works largely ignore the fact that the noise of quantum devices varies as time goes by. Therefore, reproducing the results on the same hardware will even become a problem. We analyze the reproducibility of quantum machine learning (QML) algorithms based on daily model training and execution data collection. Our analysis shows a correlation between our QML models’ test accuracy and quantum computer hardware’s calibration features. We also demonstrate that noisy simulators for quantum computers are not a reliable tool for quantum machine learning applications.

Senapati, Priyabrata↗

Sensitive Resources Assessment and Forest Analysis for The SSP-2A Parcel and Proposed Oak Ridge Enhanced Technology and Training Center (ORETTC), Oak Ridge, Tennessee

This report summarizes current knowledge of natural and cultural resources associated with potential land use changes within an 81-acre (32.8-hectare) parcel, termed SSP-2A, on the US Department of Energy’s (DOE’s) Oak Ridge Reservation (ORR) in Oak Ridge, Tennessee (Figure 1). A primary goal for the work presented here was to evaluate potential impacts to sensitive resources within the SSP-2A parcel that might result from land disturbance and construction of the Oak Ridge Enhanced Technology and Training Center (ORETTC). In addition to on-the-ground surveys of the ORETTC footprint and SSP-2A parcel during summer 2020 (Figure 1), this report leverages historical (pre-1995) and contemporary (1995–present) data from additional sources such as the Tennessee Department of Environment and Conservation (TDEC). The individuals who obtained and compiled the data that are presented here are familiar with and routinely assess, manage, and research sensitive resources on the ORR. This report should facilitate more environmentally sound decisions during planning and development of the ORETTC, provide a foundation for further assessment of sensitive and cultural resources associated with the broader SSP-2A parcel (should additional actions take place), and help project managers address regulatory guidance and DOE policy on sustainable development. Those who reference this report must consider that the timing of surveys does not permit complete delineation of resources. Data deficiencies are indicated where possible. Additional surveys may be required to account for seasonal patterns of various threatened and endangered species (e.g., bats), and additional assessment will be required if activities extend beyond the ORETTC site (Figure 2).

54 ENVIRONMENTAL SCIENCES↗

Augmented Reality Technologies for Radiation Safety Training: A Systematic Review of Sensor Integration and Visualization Approaches

This paper presents a comprehensive systematic review examining the application of augmented reality (AR) and sensor technologies for visualizing ionizing radiation in virtual training environments. The review methodology involved systematic identification and analysis of the relevant literature based on predetermined criteria including publication type, year of publication, application domain, and technological approach. The literature search encompassed publications from 2011 to 2021 across four major academic databases: Web of Science, Google Scholar, IEEE Xplore, and Scopus. Through rigorous screening following PRISMA 2020 guidelines, 23 research articles met the inclusion criteria for detailed analysis. From 404 initial database records, 360 were excluded during title/abstract screening (primarily for lacking AR components, radiation focus, or training applications) and 4 during full-text assessment (all for lacking sensor integration). The findings reveal that AR-based ionizing radiation visualization has been successfully implemented across diverse domains, including nuclear facility operations, medical procedures, CERN research activities, and educational and monitoring applications. The analysis identified multiple dimensions of impact, encompassing distinct benefits, emerging opportunities, and implementation challenges associated with AR deployment for ionizing radiation training. Each of these dimensions is comprehensively examined and documented within this review. Additionally, this study identifies critical research gaps that currently limit the full potential of AR technology in supporting ionizing radiation training programs. These gaps are systematically analyzed and discussed to establish clear directions for future research endeavors in this emerging field.

61 - RADIATION PROTECTION AND DOSIMETRY↗

DESI complete calibration of the colour–redshift relation (DC3R2): results from early DESI data

We present initial results from the Dark Energy Spectroscopic Instrument (DESI) complete calibration of the colour–redshift relation (DC3R2) secondary target survey. Our analysis uses 230 k galaxies that overlap with KiDS-VIKING ugriZYJHK s photometry to calibrate the colour–redshift relation and to inform photometric redshift (photo-z) inference methods of future weak lensing surveys. Together with emission line galaxies (ELGs), luminous red galaxies (LRGs), and the Bright Galaxy Survey (BGS) that provide samples of complementary colour, the DC3R2 targets help DESI to span 56 percent of the colour space visible to Euclid and LSST with high confidence spectroscopic redshifts. The effects of spectroscopic completeness and quality are explored, as well as systematic uncertainties introduced with the use of common Self-Organizing Maps trained on different photometry than the analysis sample. We further examine the dependence of redshift on magnitude at fixed colour, important for the use of bright galaxy spectra to calibrate redshifts in a fainter photometric galaxy sample. We find that noise in the KiDS-VIKING photometry introduces a dominant, apparent magnitude dependence of redshift at fixed colour, which indicates a need for carefully chosen deep drilling fields, and survey simulation to model this effect for future weak lensing surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Oakland University Cybersecurity Center (Final Scientific/Technical Report)

This report summarizes the outcomes of Award DE-CR0000023, “Oakland University Cybersecurity Center,” a 31-month project funded by the U.S. Department of Energy Office of Cybersecurity, Energy Security, and Emergency Response (CESER). The project addressed cybersecurity risks facing small and medium-sized manufacturers (SMMs) transitioning to Industry 4.0. The project integrated customer discovery, applied research, and cybersecurity training development. A total of 51 cybersecurity assessments identified significant gaps in baseline practices, incident response, and workforce capability. Research efforts produced a scalable mitigation framework tailored to SMM environments, and workforce analysis identified persistent talent gaps. Eight cybersecurity training modules were developed and deployed via Oakland University’s Professional and Continuing Education (PACE) platform. All objectives were completed, with 98.93% federal budget utilization and cost share exceeding requirements. The project establishes a scalable model for strengthening cybersecurity resilience and workforce capacity across U.S. manufacturing supply chains.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Accurate Thermochemistry of Complex Lignin Structures via Density Functional Theory, Group Additivity, and Machine Learning

A molecular-level understanding of lignin structures and bond dissociation energies could facilitate depolymerization technologies. Still, this information is currently limited due to the lack of databases and the simplification of surrogate models. Here, substitution effects on seven common linkages in lignin polymers are systematically investigated. An automated reaction network generator is employed to create a database of structures. A new group additivity (GA) model based on principal component analysis (PCA) descriptors is introduced and trained on gas-phase density functional theory data of 4100 species at the M06-2X/6-311++G(d,p) level. Hydrogen bonds, local steric, and nonaromatic ring contributions are also incorporated. Lastly, we improve the accuracy of the group additivity model to reach the G4 theory by computing a data set of 770 species at this level and using a data fusion approach.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Visual feedback and guided balance training in an immersive virtual reality environment for lower extremity rehabilitation

Balance training is essential for physical rehabilitation procedures, as it can improve functional mobility and enhance cognitive coordination. However, conventional balance training methods may have limitations in terms of motivation, real-time objective feedback, and personalization, which a virtual reality (VR) setup may provide a better alternative. In this work, we present an immersive VR training environment for lower extremity balance rehabilitation with real-time guidance and feedback. The VR training environment immerses the user in a 3D ice rink model where a virtual coach (agent) leads them through a series of balance poses, and the user controls a trainee avatar with their own movements. Here we developed two coaching styles: positive-reinforcement and autonomous-supportive, and two viewpoints of the trainee avatar: first-person and third-person. The proposed environment was evaluated in a user study with healthy, non-clinical participants (n = 16, 24.4 ± 5.7 years old, 9 females). Our results show that participants showed stronger performance in the positive-reinforcement style compared to the autonomous-supportive style. Additionally, in the third-person viewpoint, the participants exhibited more stability in the positive-reinforcement style compared to the autonomous-supportive style. For viewpoint, participants exhibited stronger performance in the first-person viewpoint compared to third-person in the autonomous-supportive style, while they were comparable in the positive-reinforcement style. We observed no significant effects on the foot height and number of mistakes. Furthermore, we report the analysis of user performance with balance training poses and subjective measures based on questionnaires to assess the user experience, usability, and task load. The proposed VR balance training could offer an interactive, adaptive, and engaging environment and open new potential research directions for lower extremity rehabilitation.

97 MATHEMATICS AND COMPUTING↗

Cyber-Informed Engineering (CIE) Workbook: End-of-Train (EoT) / Head-of-Train (HoT) Communications

This workbook presents a vulnerability (CVE-2025-1727 ) found in train applications and guides a digital risk assessment and mitigation analysis and application of Cyber-Informed Engineering principles to mitigate the potential consequences and ultimately the hazard through the engineering discipline because of exploiting this vulnerability. Workshop participants are encouraged to use the workbook to capture insights and lessons learned. The workbook guides the participant to: • Understand the HE communication vulnerability • Map digital threats to physical consequences • Use bowtie analysis to illustrate both “security” and “engineering” barriers • Apply CIE principles to ensure that even if communications are compromised, the physical engineered system still behaves safely. • Produce an actionable set of engineered and infosec controls for implementation

42 - ENGINEERING↗

Evaluation of pre-training large language models on leadership-class supercomputers

Large language models (LLMs) have arisen rapidly to the center stage of artificial intelligence as the foundation models applicable to many downstream learning tasks. However, how to effectively build, train, and serve such models for many high-stake and first-principle-based scientific use cases are both of great interests and of great challenges. Moreover, pre-training LLMs with billions or even trillions of parameters can be prohibitively expensive not just for academic institutions, but also for well-funded industrial and government labs. Furthermore, the energy cost and the environmental impact of developing LLMs must be kept in mind. Here, in this work, we conduct a first-of-its-kind performance analysis to understand the time and energy cost of pre-training LLMs on the Department of Energy (DOE)’s leadership-class supercomputers. Employing state-of-the-art distributed training techniques, we evaluate the computational performance of various parallelization approaches at scale for a range of model sizes, and establish a projection model for the cost of full training. Our findings provide baseline results, best practices, and heuristics for pre-training such large models that should be valuable to HPC community at large. We also offer insights and optimization strategies for using the first exascale computing system, Frontier, to train models of the size of GPT-3 and beyond.

97 MATHEMATICS AND COMPUTING↗

SGD-Net: Efficient Model-Based Deep Learning with Theoretical Guarantees

Deep unfolding networks have recently gained popularity for solving imaging inverse problems. However, the computational and memory complexity of data-consistency layers within traditional deep unfolding networks scales with the number of measurements, limiting their applicability to large-scale imaging inverse problems. We propose SGD-Net as a new methodology for improving the efficiency of deep unfolding through stochastic approximations of the data-consistency layers. Our theoretical analysis shows that SGD-Net can be trained to approximate batch deep unfolding networks to an arbitrary precision. Our simulations on intensity diffraction tomography and sparse-view computed tomography show that SGD-Net can match the performance of the traditional batch network at a fraction of training and testing complexity.Deep unfolding networks have recently gained popularity for solving imaging inverse problems. However, the computational and memory complexity of data-consistency layers within traditional deep unfolding networks scales with the number of measurements, limiting their applicability to large-scale imaging inverse problems. We propose SGD-Net as a new methodology for improving the efficiency of deep unfolding through stochastic approximations of the data-consistency layers. Our theoretical analysis shows that SGD-Net can be trained to approximate batch deep unfolding networks to an arbitrary precision. Our simulations on intensity diffraction tomography and sparse-view computed tomography show that SGD-Net can match the performance of the traditional batch network at a fraction of training and testing complexity.

97 MATHEMATICS AND COMPUTING↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

Python codes for the paper - "Trade-offs in the latent representation of microstructure evolution"

SAND2024-00946O Python codes used in "Trade-offs in the latent representation of microstructure evolution," a manuscript accepted for publication in Acta Materialia, are part of a repository. The code was developed to perform analysis of microstructure evolution. The repository consists of two main directories: models, which train and test models such as autoencoders and diffusion maps, and analysis, which analyzes microstructures. Source code is used to perform dimensionality reduction of microstructure data for analysis of its evolution in time. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dingreville, Remi↗

Rapid Inverse Parameter Inference Using Physics-Informed Neural Network

As Li-ion batteries become more essential in today's economy, tools need to be developed to accurately and rapidly diagnose a battery's internal state-of-health. Using a Li-ion battery's (high-rate) voltage response, it is proposed to determine a battery's internal state through Bayesian calibration. However, Bayesian calibration is notoriously slow and requires thousands of model runs. To accelerate parameter inference using Bayesian calibration, a surrogate model is developed to replace the underlying physics-based Li-ion model. Developing a surrogate model for rapid Bayesian calibration analysis is discussed for both the single particle model (SPM) and the pseudo two-dimensional (P2D) model. Surrogate models are constructed using physics-informed neural networks (PINNs) that encode the influence of internal properties on observed voltage responses. In practice, a neural network can be trained by: 1) using simulation results of the physics-based model (i.e., a data-loss approach); 2) using the residuals of the governing equations themselves (i.e., a physics-loss approach); or 3) using a combination of simulation results and governing equation residuals. In the present work, PINNs are developed using a variety of training losses and neural network architectures. In this analysis, it is shown that a PINN surrogate model can be reliably trained with only physics-informed loss. However, using a coupled data-informed and physics-loss approach produced the most accurate PINNs.

Bayesian calibration↗