Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high performance analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Early Exploration of a Flexible Framework for Efficient Quantum Linear Solvers in Power Systems

The rapid integration of renewable energy resources presents formidable challenges in managing power grids. While advanced computing and machine learning techniques offer some solutions for accelerating grid modeling and simulation, there remain complex problems that classical computers cannot effectively address. Quantum computing, a promising technology, has the potential to fundamentally transform how we manage power systems, especially in scenarios with a higher proportion of renewable energy sources. One critical aspect is solving linear systems of equations, crucial for power system applications like power flow analysis, for which the Harrow-Hassidim-Lloyd (HHL) algorithm is a well-known quantum solution. However, HHL quantum circuits often exhibit excessive depth, making them impractical for current Noisy-Intermediate-Scale-Quantum (NISQ) devices. In this paper, we introduce a versatile framework, powered by NWQSim, that bridges the gap between power system applications and quantum linear solvers available in Qiskit. This framework empowers researchers to efficiently explore power system applications using quantum linear solvers. Through innovative gate fusion strategies, reduced circuit depth, and GPU acceleration, our simulator significantly enhances resource efficiency. Power flow case studies have demonstrated up to a eight-fold speedup compared to Qiskit Aer, all while maintaining comparable levels of accuracy.

quantum computing, Harrow-Hassidim-Lloyd, high-per↗

Towards Interactive, Reproducible Analytics at Scale on HPC Systems

The growth in scientific data volumes has resulted in a need to scale up processing and analysis pipelines using High Performance Computing (HPC) systems. These workflows need interactive, reproducible analytics at scale. The Jupyter platform provides core capabilities for interactivity but was not designed for HPC systems. In this paper, we outline our efforts that bring together core technologies based on the Jupyter Platform to create interactive, reproducible analytics at scale on HPC systems. Our work is grounded in a real world science use case-applying geophysical simulations and inversions for imaging the subsurface. Our core platform addresses three key areas of the scientific analysis workflow-reproducibility, scalability, and interactivity. We describe our implemention of a system, using Binder, Science Capsule, and Dask software. We demonstrate the use of this software to run our use case and interactively visualize real-Time streams of HDF5 data.

containers↗

Sandia’s Liquid-Cooled Data Center Boosts Efficiency and Resiliency

The Federal Energy Management Program (FEMP) encourages federal agencies and organizations to improve data center energy efficiency, which can offer tremendous opportunities for energy and cost savings. In this success story, a novel liquid cooling system provides reliable, resilient, and energy-efficient cooling for high performance computing (HPC) systems at Sandia National Laboratories.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MOSIQS: Persistent Memory Object Storage With Metadata Indexing and Querying for Scientific Computing

Scientific applications often require high-bandwidth shared storage to perform joint simulations and collaborative data analytics. Shared memory pools provide a chance to satisfy such needs. Recently, a high-speed network such as Gen-Z utilizing persistent memory (PM) offers an opportunity to create a shared memory pool connected to compute nodes. However, there are several challenges to use scientific applications on the shared memory pool directly such as scalability, failure-atomicity, and lack of scientific metadata-based search and query. In this paper, we propose MOSIQS, a persistent memory object storage framework with metadata indexing and querying for scientific computing. We design MOSIQS based on the key idea that memory objects on PM pool can live beyond the application lifetime and can become the sharing currency for applications and scientists. MOSIQS provides an aggregate memory pool atop an array of persistent memory devices to store and access memory objects to accelerate scientific computing. MOSIQS uses a lightweight persistent memory key-value store to manage the metadata of memory objects, which enables memory object sharing. To facilitate metadata search and query over millions of memory objects resident on memory pool, we introduce Group Split and Merge (GSM), a novel persistent index data structure designed primarily for scientific datasets. GSM splits and merges dynamically to minimize the query search space and maintains low query processing time while overcoming the index storage overhead. MOSIQS is implemented on top of PMDK. We evaluate the proposed approach on many-core server with an array of real PM devices. Experimental results show that MOSIQS gains a 100% write performance improvement and executes multi-attribute queries efficiently with 2.7× less index storage overhead offering significant potential to speed up scientific computing applications.

97 MATHEMATICS AND COMPUTING↗

Improving Progressive Retrieval for HPC Scientific Data using Deep Neural Network

As the disparity between compute and I/O on high-performance computing systems has continued to widen, it has become increasingly difficult to perform post-hoc data analytics on full-resolution scientific simulation data due to the high I/O cost. Error-bounded data decomposition and progressive data retrieval framework has recently been developed to address such a challenge by performing data decomposition before storage and reading only part of the decomposed data when necessary. However, the performance of the progressive retrieval framework has been suffering from the over-pessimistic error control theory, such that the achieved maximum error of recomposed data is significantly lower than the required error. Therefore, more data than required is fetched for recomposition, incurring additional I/O overhead. In order to tackle this issue, we propose a DNN-based progressive retrieval framework that can better identify the minimum amount of data to be retrieved. Our contributions are as follows: 1) We provide an in-depth investigation of the recently developed progressive retrieval framework; 2) We propose two designs of prediction models (named D-MGARD and E-MGARD) to estimate the amount of retrieved data size based on error bounds. 3) We evaluate our proposed solutions using scientific datasets generated by real-world simulations from two domains. Evaluation results demonstrate the effectiveness of our solution in accurately predicting the amount of retrieval data size, as well as the advantages of our solution over the traditional approach to reducing the I/O overhead. Based on our evaluation, our solution is shown to read significantly less data (5% - 40% with D-MGARD, 20% - 80% with E-MGARD).

Wang, Jinzhen↗

Hamilton: Flexible, Open Source $10 Wireless Sensor System for Energy Efficient Building Operation

Sensors for improving building performance are rapidly populating the market, driven in part by the drive to reduce greenhouse gas emissions resulting from energy production as well as improve the interior environment for healthy and more productive spaces. UC Berkeley has led wireless sensor development over the past 25 years (e.g., Telos mote), with the Hamilton (named after Alexander Hamilton on the US $10 bill) as the most recent. The Hamilton sensor was designed as a low-cost high-performance sensor that is modular and interoperable. The objective of the Hamilton project was to create, evaluate and establish the technological foundations for secure and easy to deploy building energy efficiency applications utilizing pervasive, low-cost wireless sensors integrated with traditional Building Management Systems (BMS), consumer-sector building components, and powerful data analytics. The project included iterative hardware design, incorporating a high-performance database (BTrDb, http://btrdb.io/), creating and iterating the development of secure data middleware (BOSSwave, WAVE/WAVEMQ), working with and pushing the development of an open-source tiny operating system RiotOS, and implementing and improving protocols such as Thread/OpenThread and TCP/IP. The hardware benefited from careful design to drive down the cost; the design included a System-on-a-Chip (SoC), chip antenna, single crystal and five passive components. Careful design of the operating system created a low-power design to enable a long life with small batteries. The hardware included several sensors: temperature, radiant temperature, relative humidity, magnetometer, accelerometer, and light, with an optional occupancy (Passive InfraRed) sensor. The project was the basis of several applications, both internal to the research team and other researchers and professionals at other institutions. Several applications used the sensor hardware as the basis for other complex devices. Other applications used the sensors to improve building performance through interoperating with the building Heating Ventilation and Air-Conditioning (HVAC) system, such as using occupancy and/or distributed temperature sensing to reduce HVAC zone energy while still providing thermal comfort and to reduce peak loads in small commercial buildings. We demonstrated cloud-based energy analytics, implemented a schedule and a Model Predictive Controller in a small commercial building to optimize HVAC energy, occupancy and electricity price. Initial integration of these technological innovations was performed through the creation of execution containers containing the WAVE agent and various driver, proxy, or building system function logic. The research added to the understanding of efficient sensor hardware, secure middleware, time-series data management (high performance database), efficient communication protocols, and interoperating with applications and building systems. The project showed the technical effectiveness and economic feasibility of creating a low-cost, modular, and easy-to-deploy sensor. Through conversations with multiple end users, the research team discovered that many customers wanted data management and services in addition to the sensors. HamiltonIOT developed packages of sensors, border router, and data services to provide a seamless “plug-and-play” sensor deployment. Some customers were willing to pay for higher quality sensors (such as light); some customers wanted a robust enclosure (waterproof).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

STREAM: A Scalable Federated HPC Telemetry Platform

Obtaining and analyzing high performance computing (HPC) telemetry in real time is a complex task that can impact algo- rithmic performance, operating costs, and ultimately scientific outcomes. If your organization operates multiple HPC systems, filesystems, and clusters, telemetry streams can be synthesized in order to ease operational and analytics burden. In order to collect this telemetry, the Oak Ridge Leadership Computing Facility (OLCF) has deployed STREAM (Streaming Telemetry for Resource Events, Analytics, and Monitoring), which is a distributed and high-performance message bus based on Apache Kafka. STREAM collects center-wide performance information and must interface with many sources, including five HPE deployed supercomputers, each with their own Kafka cluster which is managed by HPCM. OLCF Supercomputers and their attached scratch filesystems currently send more than 300 million messages to over 200 topics producing around 1.3 Terabytes per day of telemetry data to STREAM. This paper describes the architectural principles that enable STREAM to be both resilient and highly performant while supporting multiple upstream Kafka clusters and other data sources. It also discusses the design challenges and decisions faced in adapting our existing system- monitoring infrastructure to support the first Exascale computing platform.

Adamson, Ryan↗

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

A parallel, distributed memory implementation of the adaptive sampling configuration interaction method

The many-body simulation of quantum systems is an active field of research that involves several different methods targeting various computing platforms. Many methods commonly employed, particularly coupled cluster methods, have been adapted to leverage the latest advances in modern high-performance computing. Selected configuration interaction (sCI) methods have seen extensive usage and development in recent years. However, the development of sCI methods targeting massively parallel resources has been explored only in a few research works. Here, we present a parallel, distributed memory implementation of the adaptive sampling configuration interaction approach (ASCI) for sCI. In particular, we will address the key concerns pertaining to the parallelization of the determinant search and selection, Hamiltonian formation, and the variational eigenvalue calculation for the ASCI method. Load balancing in the search step is achieved through the application of memory-efficient determinant constraints originally developed for the ASCI-PT2 method. The presented benchmarks demonstrate near optimal speedup for ASCI calculations of Cr 2 (24e, 30o) with 10 6 , 10 7 , and 3 × 10 8 variational determinants on up to 16 384 CPUs. Importantly, to the best of the authors’ knowledge, this is the largest variational ASCI calculation to date.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

The AI Engine (AIE) architecture, available in systems from mobile SoCs to server-class FPGAs, aims to efficiently execute AI/ML tasks through a two-dimensional array of compute tiles. Previous work on AIEs has explored different approaches to mapping computation across spatial arrays, but the compute kernel running on each tile has not been the focus. Additionally, the AIE-ML architecture introduces memory tiles and omits programmable logic, requiring new approaches to staging and moving data throughout the array. In this work we update analytical models developed for CPUs to produce the design of high performance kernels while introducing new model considerations such as memory structure, throughput, and latency as required by the AIE hardware. We evaluate our models by developing AIE-ML kernels for matrix multiplication in low-precision data types showing performance up to 95% of compute peak for the kernel when data resides in local memory and above 90% of compute peak when data resides in main memory.

Binder, Elliott D. [Carnegie Mellon University, Pi↗

Graph Analytics on Jellyfish topology

Because large unstructured datasets is important for many science domains, distributed graph analytics is critical to many scientists. Unfortunately, obtaining scaling and performance for irregular communication is challenging because contemporary network interconnects are primarily designed to maximize bandwidths of fixed-neighborhoods large-message exchanges (e.g., stencils). Although there is no consensus on the “best” network topologies for irregular communication, unstructured graph-based interconnects can be more suitable. We analyze three popular graph workloads – clustering, pattern enumeration, and traversal — on comparable networks (in terms of resources and costs) constructed from Jellyfish Random Regular, Dragonfly and Fat tree topologies, varying the routing algorithms. Using packet-level simulations, we demonstrate up to 60% improvement in communication time with Jellyfish due to diversity of the short paths between arbitrary endpoints, which can reduce overall network stalls and congestion.

Graph Analytics, network topology, interconnect, H↗

4D Printing via an Unconventional Fused Deposition Modeling Route to High-Performance Thermosets

An unprecedented four-dimensional (4D) printing process allowing high-performance and shape memory thermoset to be printed, for the first time, by fused deposition modeling (FDM) with isotropic properties has been achieved. Here, bisphenol A-based epoxy and benzoxazine were formulated to a low-temperature thermoplastic and high-temperature thermoset resin, which is melt-extrudable and can be postcured into covalently cross-linked material. Carbon nanotube (CNT) was added in the resin to work as both mechanical enhancement filler and rheology modifier to prevent shape deformation during postcuring process. The cross-layer reaction fuses individual layers into an integrity, thus eliminating layer delamination induced by FDM, offering isotropic mechanical properties regardless of the printing orientations. The highly cross-linked network provides outstanding mechanical strength and superb thermal stability. The excellent shape memory performance with fast recovery rate and large recovery degree is also obtained in the three-dimensional (3D) printed composites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Kinetics and Optimization of the Lysine–Isopeptide Bond Forming Sortase Enzyme from Corynebacterium diphtheriae

Site-specifically modified protein bioconjugates have important applications in biology, chemistry, and medicine. Functionalizing specific protein side chains with enzymes using mild reaction conditions is of significant interest, but remains challenging. Recently, the lysine–isopeptide bond forming activity of the sortase enzyme that builds surface pili in Corynebacterium diphtheriae ( Cd SrtA) has been reconstituted in vitro. A mutationally activated form of Cd SrtA was shown to be a promising bioconjugating enzyme that can attach Leu-Pro-Leu-Thr-Gly peptide fluorophores to a specific lysine residue within the N-terminal domain of the SpaA protein ( N SpaA), enabling the labeling of target proteins that are fused to N SpaA. Here we present a detailed analysis of the Cd SrtA catalyzed protein labeling reaction. We show that the first step in catalysis is rate limiting, which is the formation of the Cd SrtA-peptide thioacyl intermediate that subsequently reacts with a lysine ε-amine in N SpaA. This intermediate is surprisingly stable, limiting spurious proteolysis of the peptide substrate. We report the discovery of a new enzyme variant ( Cd SrtA Δ ) that has significantly improved transpeptidation activity, because it completely lacks an inhibitory polypeptide appendage (“lid”) that normally masks the active site. We show that the presence of the lid primarily impairs formation of the thioacyl intermediate and not the recognition of the N SpaA substrate. Quantitative measurements reveal that Cd SrtA Δ generates its cross-linked product with a catalytic turnover number of 1.4 ± 0.004 h –1 and that it has apparent K M values of 0.16 ± 0.04 and 1.6 ± 0.3 mM for its N SpaA and peptide substrates, respectively. Cd SrtA Δ is 7-fold more active than previously studied variants, labeling >90% of N SpaA with peptide within 6 h. The results of this study further improve the utility of Cd SrtA as a protein labeling tool and provide insight into the enzyme catalyzed reaction that underpins protein labeling and pilus biogenesis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Understanding the Impact of Data Staging for Coupled Scientific Workflows

We report the rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows.

97 MATHEMATICS AND COMPUTING↗

Expediting field-effect transistor chemical sensor design with neuromorphic spiking graph neural networks

Improving the sensitive and selective detection of analytes in a variety of applications requires accelerating the rational design of field-effect transistor (FET) chemical sensors. Achieving high-performance detection relies on identifying optimal probe materials that can effectively interact with target analytes, a process traditionally driven by chemical intuition and time-consuming trial-and-error methods. To address the difficulties in probe screening for FET sensor development, this work presents a methodology that combines neuromorphic machine learning (ML) architectures, specifically a hybrid spiking graph neural network (SGNN), with an enriched dataset of physicochemical properties through semi-automated data extraction using large language models. Achieving a classification accuracy of 0.89 in predicting sensor sensitivity categories, the SGNN model outperformed traditional ML techniques by leveraging its ability to capture both global physicochemical properties and sparse topological features through a hybrid modeling framework. Next-generation sensor design was informed by the actionable insights into the connections between material properties and sensing performance offered by the SGNN framework. Through virtual screening for the detection of per- and polyfluoroalkyl substances (PFAS) as a use case, the effectiveness of the SGNN model was further validated. Density functional theory simulations confirmed graphene as a promising active material for PFAS detection as suggested by the SGNN framework. By bridging gaps in predictive modeling and data availability, this integrated approach provides a strong foundation for accelerating advancements in FET sensor design and innovation.

Ferreira, Rodrigo Pires [Univ. of Chicago, IL (Uni↗

Correlation energy of the uniform electron gas determined by ground-state conditional probability density functional theory

Conditional-probability density functional theory (CP-DFT) is a formally exact method for finding correlation energies from Kohn-Sham DFT without evaluating an explicit energy functional. We present details on how to generate accurate exchange-correlation energies for the ground-state uniform gas. We also use the exchange hole in a CP antiparallel spin calculation to extract the high-density limit. We give a highly accurate analytic solution to the Thomas-Fermi model for this problem, showing its performance relative to Kohn-Sham and may be useful at high temperatures. We explore several approximations to the CP potential. Furthermore, results are compared to accurate parameterizations for both exchange-correlation energies and holes.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Extreme-ultraviolet spatiotemporal vortices via high harmonic generation

Spatiotemporal optical vortices (STOVs) are space–time structured light pulses with a unique topology that couples spatial and temporal domains and carry transverse orbital angular momentum (OAM). Up to now, their generation has been limited to the visible and infrared regions of the spectrum. During the last decade, it was shown that through the process of high-order harmonic generation, it is possible to upconvert spatial optical vortices that carry longitudinal OAM from the near-infrared into the extreme-ultraviolet (EUV), thereby producing vortices with distinct femtosecond and attosecond structure. In this work, we demonstrate theoretically and experimentally the generation of EUV spatiotemporal and spatiospectral vortices using near-infrared STOV driving laser pulses. Here, we use analytical expressions for focused STOVs to perform macroscopic calculations of high-order harmonic generation that are directly compared to the experimental results. As STOV beams are not eigenmodes of propagation, we characterize the highly charged EUV STOVs in both the near and far fields to show that they represent conjugated spatiotemporal and spatiospectral vortex pairs. Our work provides high-frequency light beams topologically coupled at the nanometre/attosecond scales domains with transverse OAM that could be suitable to explore electronic dynamics in magnetic materials, chiral media and nanostructures.

High-harmonic generation↗

A population data-driven workflow for COVID-19 modeling and learning

CityCOVID is a detailed agent-based model that represents the behaviors and social interactions of 2.7 million residents of Chicago as they move between and colocate in 1.2 million distinct places, including households, schools, workplaces, and hospitals, as determined by individual hourly activity schedules and dynamic behaviors such as isolating because of symptom onset. Disease progression dynamics incorporated within each agent track transitions between possible COVID-19 disease states, based on heterogeneous agent attributes, exposure through colocation, and effects of protective behaviors of individuals on viral transmissibility. Throughout the COVID-19 epidemic, CityCOVID model outputs have been provided to city, county, and state stakeholders in response to evolving decision-making priorities, while incorporating emerging information on SARS-CoV-2 epidemiology. Here we demonstrate our efforts in integrating our high-performance epidemiological simulation model with large-scale machine learning to develop a generalizable, flexible, and performant analytical platform for planning and crisis response.

97 MATHEMATICS AND COMPUTING↗