Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AI efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Cybersecurity Considerations for Emerging Energy Technologies

AI, cloud computing, post-quantum cryptography, zero-trust architectures, microgrids, and virtual power plants. What do these things have in common? They are all emerging technologies in the clean energy space that will be a critical part of grid modernization efforts. As we work towards clean energy and decarbonization targets, these technologies, developed to solve real-world problems, will help us reach goals and achieve new efficiencies as the paradigm of grid operation shifts. However, there are growing concerns about the cybersecurity risks associated with these trending topics as they are used in critical infrastructure applications. This talk will cover gaps, challenges, and opportunities for the secure implementation of grid modernization solutions and novel energy applications of state-of-the-art networking and communications. Proactive risk mitigation strategies, including the application of cyber-informed engineering, will be discussed. Practical applications of these techniques will help provide countermeasures to the impact of cyberattacks on critical infrastructure technologies in a new, digitized grid landscape.

14 SOLAR ENERGY↗

Performance results of cooperating expert systems in a distributed real-time monitoring system

There are numerous definitions for real-time systems, the most stringent of which involve guaranteeing correct system response within a domain-dependent or situationally defined period of time. For applications such as diagnosis, in which the time required to produce a solution can be non-deterministic, this requirement poses a unique set of challenges in dynamic modification of solution strategy that conforms with maximum possible latencies. However, another definition of real time is relevant in the case of monitoring systems where failure to supply a response in the proper (and often infinitesimal) amount of time allowed does not make the solution less useful (or, in the extreme example of a monitoring system responsible for detecting and deflecting enemy missiles, completely irrelevant). This more casual definition involves responding to data at the same rate at which it is produced, and is more appropriate for monitoring applications with softer real-time constraints, such as interplanetary exploration, which results in massive quantities of data transmitted at the speed of light for a number of hours before it even reaches the monitoring system. The latter definition of real time has been applied to the MARVEL system for automated monitoring and diagnosis of spacecraft telemetry. An early version of this system has been in continuous operational use since it was first deployed in 1989 for the Voyager encounter with Neptune. This system remained under incremental development until 1991 and has been under routine maintenance in operations since then, while continuing to serve as an artificial intelligence (AI) testbed in the laboratory. The system architecture has been designed to facilitate concurrent and cooperative processing by multiple diagnostic expert systems in a hierarchical organization. The diagnostic modules adhere to concepts of data-driven reasoning, constrained but complete nonoverlapping domains, metaknowledge of global consequences of anomalous data, hierarchical reporting of problems that extend beyond a single domain, and shared responsibility for problems that overlap domains. The system enables efficient diagnosis of complex system failures in real-time environments with high data volumes and moderate failure rates, as indicated by extensive performance measurements.

Schwuttke, U. M.↗

NASA POWER: Providing Analysis-Ready, Cloud-Optimized Data for AI /ML Training and Applications in Earth Science

As global demand for sustainable development grows, the integration of Earth Observation (EO) data into decision making frameworks has become a primary objective for the scientific community. The NASA Prediction of Worldwide Energy Resources (POWER) project serves as a bridge between NASA EO data and the specialized needs of the renewable energy, sustainable infrastructure and agroclimatology communities. In this poster presentation we will present an overview of POWER data products and services along with its use in diverse research to decision-making workflows. By providing over 40 years of high-resolution historical, hourly and daily solar and meteorological data, POWER transforms satellite observations and global model reanalysis into actionable, Analysis-Ready Dataset (ARD). Currently, the project delivers over 250 industry-friendly parameters to the users from different NASA datasets like CERES SYN1Deg, MERRA-2, and IMERG alongside downscaled CMIP6 climate model data, fulfilling over 16 million requests from 50,000 unique users monthly. To ensure data quality and traceability, these parameters are rigorously validated against the ground-based observations from the Baseline Surface Radiation Network (BSRN) and the Global Surface Summary of the Day (GSOD) – these results will be discussed in the presentation. A newly introduced web-based PaRameter Uncertainty ViEwer (PRUVE) tool will be presented that provides an online validation platform to the users that benchmarks satellite-based and assimilation data products against these surface measurements. To reduce technical barriers to data adoption, POWER data is accessible through RESTful APIs, ESRI ArcGIS Image Services, a web-based Data Access Viewer tool, allowing users to visualize, validate and apply the dataset. For efficient data delivery POWER data is cloud-optimized into Zarr datastore accessible through NASA managed Amazon S3 ensures high-performance allowing users to integrate EO directly into operational pipelines. These customized services will be presented. Use cases from application will be presented from the energy sector - such as for design of generation systems, performance monitoring of solar power plants, in infrastructure sector- optimizing building energy efficiency and thermal comfort, in agriculture – such as driving crop simulation and yield forecasting models to enable climate resilient farming. Furthermore, the shift toward machine learning (ML) in EO research that has positioned POWER as a key provider for training datasets which will be discussed. Use-cases will be presented to showcase how NASA data is enabling the development of predictive tools for climate variability and resource management. The poster will present POWER’s future plans including technology development to enhance data traceability and reproducibility and improving I/O performance to support the rapid integration of new EO products, ensuring that POWER remains a robust scalable backend for the evolving landscape of AI-driven Earth Science. Additionally, POWER is developing an AI Agent and an MCP-Server to enable industry AI-Agentic workflows.

Neha Khadka↗

Chemical looping based ammonia production - A promising pathway for production of the noncarbon fuel

Ammonia, primarily made with Haber–Bosch process developed in 1909 and winning two Nobel prizes, is a promising noncarbon fuel for preventing global warming of 1.5 °C above pre-industrial levels. However, the undesired characteristics of the process, including high carbon footprint, necessitate alternative ammonia synthesis methods, and among them is chemical looping ammonia production (CLAP) that uses nitrogen carrier materials and operates at atmospheric pressure with high product selectivity and energy efficiency. To date, neither a systematic review nor a perspective in nitrogen carriers and CLAP has been reported in the critical area. So, this work not only assesses the previous results of CLAP but also provides perspectives towards the future of CLAP. It classifies, characterizes, and holistically analyzes the fundamentally different CLAP pathways and discusses the ways of further improving the CLAP performance with the assistance of plasma technology and artificial intelligence (AI).

30 DIRECT ENERGY CONVERSION↗

Baseflow Identification via Explainable AI With Kolmogorov‐Arnold Networks

Abstract Hydrological models often involve constitutive laws that may not be optimal in every application. We propose to replace such laws with the Kolmogorov‐Arnold networks (KANs), a class of neural networks designed to identify symbolic expressions. We demonstrate KAN's potential on the problem of baseflow identification, a notoriously challenging task plagued by significant uncertainty. KAN‐derived functional dependencies of the baseflow components on the aridity index outperform their original counterparts; they demonstrate that water availability, rather than potential evapotranspiration, drives baseflow by constraining actual evapotranspiration under arid conditions. On a test set, they increase the Nash‐Sutcliffe efficiency (NSE) by 65%, decrease the root mean squared error by 29%, and increase the Kling‐Gupta efficiency by 34%. This superior performance is achieved while reducing the number of fitting parameters from three to two. Next, we use data from 378 catchments across the continental United States to refine the water‐balance equation at the mean‐annual scale. The KAN‐derived equations based on the refined water balance outperform both the current aridity index model, with up to a 105% increase in NSE, and the KAN‐derived equations based on the original water balance. While the performance of our model and tree‐based machine learning methods is similar, KANs offer the advantage of simplicity and transparency and require no specific software or computational tools. This case study focuses on the aridity index formulation, but the approach is flexible and transferable to other hydrological processes. Plain Language Summary Equations used in hydrologic model are often suboptimal, resulting in reduced prediction accuracy and efficiency. We implemented Kolmogorov‐Arnold networks (KAN), a machine learning algorithm for deriving symbolic formulations, to estimate groundwater recharge and showed that it outperforms an existing state‐of‐the‐art semi‐empirical formulation. In hydrology, Nash‐Sutcliffe efficiency (NSE), root mean squared error (RMSE), and Kling‐Gupta efficiency (KGE) are commonly used to evaluate model performance. Higher NSE and KGE values indicate better performance, while lower RMSE values are preferable. Our results show that NSE increased by 71%, RMSE decreased by 32%, and KGE improved by 25%. In addition, KAN identifies an optimal functional form and can be used to derive new analytical formulas using the prior knowledge. The KAN‐inspired equation outperformed the original formulation and reduced the fitting parameters. Furthermore, we refined the water‐balance equation at the mean‐annual scale and showed that, based on the new water‐balance equation, KAN can derive new formulations that are superior to the original aridity index formulations (up to 105% increase in NSE) and KAN‐derived equations based on the original water balance. These findings highlight the significant potential of KAN to advance the scientific understanding of a wide range of hydrologic processes. Key Points Kolmogorov‐Arnold networks (KANs) enhance interpretability of machine‐learned hydrological models KAN‐derived symbolic formulations outperform state‐of‐the‐art semi‐empirical aridity indices KAN‐identified functional form yields an analytical index with fewer fitting parameters and improved performance

baseflow↗

Multi-Source Machine Learning and Thermoplastics Enhanced Aerostructure Manufacturing (mTEAM)

RTX Technology Research Center (RTRC), together with Collins Aerospace (Collins) and Oak Ridge National Laboratory (ORNL) has developed an Artificial Intelligence (AI) / Machine Learning (ML) guided solution to advance the manufacturing and assembly of high performance and lightweight thermoplastic composite (TPC) aerospace products. The solution aims to lower risk, cost and lead time for induction heating based welding and consolidation processes for TPC structure. The cost and lead time of part and material specific process development for induction welding (IW) and induction consolidation will be reduced by replacing traditional empirical methods with optimization methods that merge AI/ML and physics-based process simulations and process experiments with sensing and controls. TPC-IW process development is empirical in nature, and uncertainties in material & process behavior exist near & far from the induction coil. Physics-based simulations can be leveraged directly for process optimization but can be too computationally expensive to run in high fidelity and real time to do robust process optimization. The key impact of successful TPC induction consolidation and welding is cost & lead time reduction for part & material specific consolidation and welding recipes. This is an enabler for more rapid deployment of TPC structures via joining assembly, which can reduce energy & cost intensive usage of autoclaves & ovens. The solution aimed to advance the U.S. Department of Energy’s interests in using thermoplastics and automation in composite manufacturing for improvement of products for existing markets via increased production speeds, reduced costs, and lowered use of energy. Welded TPC structures can offer significant weight & energy savings for high-value commercial aerospace & industrial applications compared to metal & thermoset composite structures assembled by mechanical fastening and/or adhesive bonding. The project was organized into two Budget Periods. Budget Period 1 (BP1) was 15 months and its goal was to perform ML process optimization framework development & deployment on lab-coupon aerostructure components. A Go/No-Go Review was performed at the end of BP1 to verify fulfilment of key tasks & milestones to justify a Go Decision to move into the next Budget Period. Budget Period 2 (BP2) was 12 months and its goal was the deployment of the ML framework for ML process optimization of pilot industrial scale aerostructure components. The overall project aim was to develop & demonstrate ML-enhanced modeling framework that learns process-property mapping from multiple data sources at different fidelities. During BP1, the team accomplished key tasks & milestones to demonstrate the concept of multi-source ML for TPC aerostructure consolidation and assembly. First, the team completed documentation of induction based TPC heating requirements including baseline metrics to compare measured results against. Next the team completed demonstration of data generation from physics-based simulations for ML surrogate model generation and demonstrated the integration of physics-based simulation data into multi-source AI/ML algorithms. In parallel, the team established the lab-coupon scale induction welding system and completed a process to label and reduce generated data from physics-based simulation and experiments for ML surrogate models to enable multi-source ML model training & testing. To complete BP1, the team integrated physics-based simulation data and experimental data into multi-source ML algorithms. This was based on the team completing ML deployment of the induction welding on a lab system at RTRC and AI/ML deployment on existing induction welding line at Collins. ORNL visited both Collins and RTRC sites to witness the TPC induction welding process. Then, ORNL designed and constructed a new version of their vision-based sensing system better adapted to acquire process signals of the TPC induction welding process for process anomaly and defect detection. In BP2, the team accomplished key tasks & milestones to scale up multi-source ML for TPC aerostructure consolidation and assembly from the lab-coupon scale to the pilot-industrial scale. In BP2, the team demonstrated real time anomaly & defect detection via experiments performed by ORNL & RTRC. The team completed ML-optimization heating trials for TPC induction consolidation at Collins, and the team confirmed pilot industrial scale experimental data from Collins was compatible with the developed ML pipeline from RTRC. The team completed sub-element scale ML process optimization demonstration at RTRC, where the team leveraged RTRC’s robotic TPC welding setup to de-risk the ML process optimization by performing ML analysis of recorded temperatures to account for complex part features. Then, the team applied its ML-derived control strategies and ML process optimization framework at Collins to the pilot-industrial scale on a demo skin-stiffener part representative of a nacelle aerostructure fan cowl section. The key innovation is the AI/ML framework enabling effective process development of high performance, lightweight, energy efficient TPCs for composite aircraft structures.

36 MATERIALS SCIENCE↗

AI‐Driven Robot Enables Synthesis‐Property Relation Prediction for Metal Halide Perovskites in Humid Atmosphere

Materials Acceleration Platforms (MAPs) – also known as self-driving laboratories– present a new paradigm for materials science and promise an order of magnitude accelerated materials discovery compared to the traditional trial-and-error approach. Metal halide perovskites (MHPs) are an emerging class of materials for optoelectronic applications but are plagued by irreproducible optoelectronic quality, particularly for films fabricated in a humid atmosphere. Here, in this work, a machine learning (ML)-guided closed-loop platform is developed with a multimodal data fusion approach to predict synthesis–property relations for the optical quality of MHP thin films in relative humidities (RHs) ranging from 5–55%. The efficiency of this approach is confirmed by the fast-dropping learning rate to 2% after experimentally sampling less than 1% of the possible 5,000+ combinations. The prediction of synthesis–property relations is done by optical and imaging characterizations. In situ photoluminescence characterization revealed the origin of thin film quality variation at different RH. These insights provide an avenue for controlling the MHP crystallization by fine-tuning the synthesis parameters and RH for a given chemistry, thus lifting the need for stringent atmosphere control. The MAP enables an accelerated screening and understanding of the synthesis design space, facilitating rational synthesis recipe choice for a wide range of materials.

AI-driven robot↗

ACTIVE

The Automated Control Testbed for Integration, Verification, and Emulation (ACTIVE) framework is a software platform designed to support the optimized operation and management of a wide range of building types. It enables the development, testing, and validation of diverse control strategies, including AI-based, rule-based, and model-based approaches. The platform facilitates a seamless transition from simulation-based evaluation of control strategies to real-world field validation and deployment. ACTIVE supports the full building management lifecycle, encompassing data acquisition and management, system monitoring, optimized control, adaptive learning services, device dispatch and coordination, as well as advanced analytics and visualization. Together, these capabilities provide an integrated environment for improving building performance, operational efficiency, reducing energy cost, and reliability.

Smith, Robert [Oak Ridge National Laboratory (ORNL↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

AI Denoising to Accelerate Detector Simulation

Detector simulation is critical to experimental HEP; however this simulation (commonly done through toolkits such as Geant4) is computationally intensive. Performance can be improved somewhat through technical optimization, but more is needed. Using machine learning (ML) to accelerate simulation is a promising field, however efforts to use generative adversarial networks (GANs) or optimized autoencoders have faced issues. Using convolutional neural networks (CNNs) for denoising has been successful in non-HEP applications such as image processing. This poster investigates the efficacy of using CNNs to denoise Geant4 simulations. This could increase the accuracy of simulations performed under settings designed to increase computational efficiency.

Franklin, Lena↗

Coupled Lattice Boltzmann Modeling Framework for Pore-Scale Fluid Flow and Reactive Transport

In this paper, we propose a modeling framework for pore-scale fluid flow and reactive transport based on a coupled lattice Boltzmann model (LBM). We develop a modeling interface to integrate the LBM modeling code parallel lattice Boltzmann solver and the PHREEQC reaction solver using multiple flow and reaction cell mapping schemes. The major advantage of the proposed workflow is the high modeling flexibility obtained by coupling the geochemical model with the LBM fluid flow model. Consequently, the model is capable of executing one or more complex reactions within desired cells while preserving the high data communication efficiency between the two codes. Meanwhile, the developed mapping mechanism enables the flow, diffusion, and reactions in complex pore-scale geometries. We validate the coupled code in a series of benchmark numerical experiments, including 2D single-phase Poiseuille flow and diffusion, 2D reactive transport with calcite dissolution, as well as surface complexation reactions. The simulation results show good agreement with analytical solutions, experimental data, and multiple other simulation codes. In addition, we design an AI-based optimization workflow and implement it on the surface complexation model to enable increased capacity of the coupled modeling framework. Compared to the manual tuning results proposed in the literature, our workflow demonstrates fast and reliable model optimization results without incorporating pre-existing domain knowledge.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Addressing Rising Energy Demand Through Innovation

The U.S. is facing a significant increase in energy demand, driven by AI advancements, the rapid expansion of data centers, manufacturing and industrial growth, and the electrification of transportation and buildings. Buildings alone account for approximately 75% of U.S. electricity consumption and 40% of total energy use. To address these challenges, NLR leverages its state-of-the-art research facilities, advanced energy modeling, hardware-in-the-loop emulation, and real-world demonstrations to provide data-driven insights that de-risk emerging energy solutions, increase efficiency and demand flexibility, optimize grid controls, and identify vulnerabilities to enhance energy security. This presentation will highlight our research ecosystem and its role in supporting a more reliable, affordable, and adaptive energy infrastructure in the face of accelerating demand.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Proceedings of the Airborne Imaging Spectrometer Data Analysis Workshop

The Airborne Imaging Spectrometer (AIS) Data Analysis Workshop was held at the Jet Propulsion Laboratory on April 8 to 10, 1985. It was attended by 92 people who heard reports on 30 investigations currently under way using AIS data that have been collected over the past two years. Written summaries of 27 of the presentations are in these Proceedings. Many of the results presented at the Workshop are preliminary because most investigators have been working with this fundamentally new type of data for only a relatively short time. Nevertheless, several conclusions can be drawn from the Workshop presentations concerning the value of imaging spectrometry to Earth remote sensing. First, work with AIS has shown that direct identification of minerals through high spectral resolution imaging is a reality for a wide range of materials and geological settings. Second, there are strong indications that high spectral resolution remote sensing will enhance the ability to map vegetation species. There are also good indications that imaging spectrometry will be useful for biochemical studies of vegetation. Finally, there are a number of new data analysis techniques under development which should lead to more efficient and complete information extraction from imaging spectrometer data. The results of the Workshop indicate that as experience is gained with this new class of data, and as new analysis methodologies are developed and applied, the value of imaging spectrometry should increase.

Vane, G.↗

Trilateral Task Force – Reliability Analysis Supporting Mission Extension/Post Mission Disposal

At the intersection of mission, technology, and place is NASA’s need to modernize for a digital-forward future. Digitalization, the process of moving toward digital business, is occurring everywhere and remains an ongoing process across the federal government.”[1] Whereas, Digital Transformation is “employing digitization/digital technologies (e.g., Artificial Intelligence (AI), mobile, cloud, data) to change a process, product, or capability so dramatically (e.g., real-time, intelligent, personalized, anywhere, anytime) that it is unrecognizable compared to its traditional form.” [2] In order to facilitate a digital transformation it is essential for NASA to understand and identify where data exists today and which data are value-needed in the future, understand where there are unfulfilled data needs that limit the advancement of NASA work, and ensure NASA efficiency through Findable, Accessible, Interoperable, and Reusable (FAIR) digital assets in the future. Therefore, NASA’s Reliability & Maintainability (R&M) Enterprise Data Sharing team is working to leverage both Digitization and Digital Transformation to achieve their vision of developing an R&M data discovery framework that enables our community, our partners, and our stakeholders with the ability to efficiently, robustly, and seamlessly access information that enables real-time knowledge and model-based, analytics driven, decision-making impacting R&M. As a result the R&M Enterprise Data Sharing team has conducted a survey of its Reliability, Maintainability, and Availability (RMA) community members to identify data existence (created or used) and where there are corresponding barriers to data acquisition and/or R&M or other issues as shown within this presentation.

Digital Transformation, Reliability Engineering↗

kynema-fmb [SWR-23-07]

Kynema-FMB (FKA: Kynema) is an open-source performance portable flexible multibody (FMB) dynamics solver designed for time-domain simulations. While originally tailored for wind turbine structural dynamics, the formulation and implementation are those of a general flexible-multidbody dynamics solver that can readily be applied to a wide range of systems. Kynema was designed with a narrow focus, namely to provide a lightweight, fast, accurate FMD solver for coupling to computational-fluid-dynamics (CFD) codes, especially the CFD codes in the Kynema suite, for fluid-structure-interaction (FSI) simulations. Kynema-FMB is equipped to model systems that can be represented as a collection of beams and rigid bodies that are connected through constraints. Degrees of freedom are defined in the inertial/global frame of reference and include displacements and rotations (formally as rotation matrices, but stored as quaternions). The underlying formulation is built on a Lie-group time integrator designed for index-3 differential-algebraic equations, which is second-order accurate in time (Bruls et al., 2012). Beam models are based on geometrically exact beam theory and are discretized as high-order spectral finite elements similar to those in BeamDyn (Wang et al., 2017). The governing equations for a FMD system like a wind turbine constitute a highly nonlinear system of constrained partial-differential equations. Kynema-FMB uses analytical Jacobians in the nonlinear-system solves in each time step. Linear systems use sparse storage and several third-party sparse-linear-system solvers are enabled. Ill conditioning of linear systems is mitigated with preconditioning described in Bottasso et al, 2008. Kynema-FMB is integrated with a simple open-source controller (ROSCO). There is an application programming interface (API) for coupling to geometry-resolved CFD (like that in Sharma et al., 2023) and actuator-force CFD (like that in Kuhn et al., 2025). In the latter, for actuator-line models, Kynema-FMB includes an internal blade-element solver that depends on user-provided lookup tables for coefficients of lift and drag, i.e., aerodynamic polars. Kynema-FMB is written in C++ and leverages Kokkos and Kokkos-Kernels (KokkosEcosystem) as its performance portability layer enabling simulations on both CPU and GPU systems. The repository is equipped with extensive automated testing at the unit and regression/system levels. The following describes the high-level development objectives conceived for Kynema: *Kynema will follow modern software development best practices, including test-driven development (TDD), version control, hierarchical automated testing, and continuous integration (CI) for a robust development environment. *The core data structures are memory efficient and enable vectorization and parallelization at multiple levels. *Data structures are data-oriented to exploit methods for accelerated computing including high utilization of chip resources (e.g., single instruction multiple data (SIMD) instruction sets) and parallelization using GP-GPUs. *The computational algorithms incorporate robust open-source libraries for mathematical operations, resource allocation, and data management. *The API design considers multiple stakeholder needs and ensure integration with existing and future ecosystems for data science, machine learning, and AI. *Kynema-FMB is written in modern C++ and leverages Kokkos as its performance-portability library with inspiration from the kynema stack.

Sprague, MichaelA.↗

Hvac: Removing I/O Bottleneck for Large-Scale Deep Learning Applications

Scientific communities are increasingly adopting deep learning (DL) models in their applications to accelerate scientific discovery processes. However, with rapid growth in the computing capabilities of HPC supercomputers, large-scale DL applications have to spend a significant portion of training time performing I/O to a parallel storage system. Previous research works have investigated optimization techniques such as prefetching and caching. Unfortunately, there exist non-trivial challenges to adopting the existing solutions on HPC supercomputers for large-scale DL training applications, which include non-performance and/or failures at extreme scale, lack of portability and generality in design, complex deployment methodology, and being limited to a specific application or dataset. To address these challenges, we propose High-Velocity AI Cache (HVAC), a distributed read-cache layer that targets and fully exploits the node-local storage or near node-local storage technology. HVAC seamlessly accelerates read I/O by aggregating node-local or near node-local storage, avoiding metadata lookups and file locking while preserving portability in the application code. We deploy and evaluate HVAC on 1,024 nodes (with over 6000 NVIDIA V100 GPUS) of the Summit supercomputer. In particular, we evaluate the scalability, efficiency, accuracy, and load distribution of HVAC compared to GPFS and XFS-on-NVMe. With four different DL applications, we observe an average 25 % performance improvement atop GPFS and 9% drop against XFS-on-NVMe, which scale linearly and are considered the performance upper bound. We envision HVAC as an important caching library for upcoming HPC supercomputers such as Frontier.

Khan, Awais↗

Certifying almost all quantum states with few single-qubit measurements

Certifying that an n -qubit state synthesized in the laboratory is close to a given target state is a fundamental task in quantum information science. However, existing rigorous protocols applicable to general target states have potentially prohibitive resource requirements in the form of either deep quantum circuits or exponentially many single-qubit measurements. Here we prove that almost all n -qubit target states, including those with exponential circuit complexity, can be certified from only O ( n 2 ) single-qubit measurements. Given access to the target state’s amplitudes, our protocol requires only O ( n 3 ) classical computation. This result is established by a technique that relates certification to the mixing time of a random walk. Our protocol has applications for benchmarking quantum systems, for optimizing quantum circuits to generate a desired target state and for learning and verifying neural networks, tensor networks and various other representations of quantum states using only single-qubit measurements. We show that such verified representations can be used to efficiently predict highly non-local properties of a synthesized state that would otherwise require an exponential number of measurements on the state. We demonstrate these applications in numerical experiments with up to 120 qubits and observe an advantage over existing methods such as cross-entropy benchmarking.

information theory and computation↗

Large Scale Caching and Streaming of Training Data for Online Deep Learning

The training of deep neural network models on large data remains a difficult problem, despite progress towards scalable techniques. In particular, there is a mismatch between the random but predetermined order in which AI flows select training samples and the streaming I/O patterns for which traditional HPC data storage (e.g., parallel file systems) are designed. In addition, as more data are obtained, it is feasible neither simply to train learning models incrementally, due to catastrophic forgetting (i.e., bias towards new samples), nor to train frequently from scratch, due to prohibitive time and/or resource constraints. In this paper, we study data management techniques that combine caching and streaming with rehearsal support in order to enable efficient access to training samples in both offline training and continual learning. We revisit state-of-art streaming approaches based on data pipelines that transparently handle prefetching, caching, shuffling, and data augmentation, and discuss the challenges and opportunities that arise when combining these methods with data-parallel training techniques. We also report on preliminary experiments that evaluate the I/O overheads involved in accessing the training samples from a parallel file system (PFS) under several concurrency scenarios, highlighting the impact of the PFS on the design of the data pipelines.

data pipelines↗