Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high performance analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Characterizing Reactor Operations from Realistic Simulated Environmental Samples: Combining High-Performance Computing and Data Analytics

Environmental sampling is a common technique employed by inspectors and facility operators in nuclear safeguards, proliferation detection, and process monitoring contexts. Interpreting measurements performed on samples or collections of samples and ensuring the information extracted is accurate and precise is difficult. To date, these analyses have relied on simulated data to enable systematic studies; however, these models are inherently limited by the fidelity of the models and the implicit spatial averaging of isotopic composition or other signatures of interest. To advance this capability, we have refined the spatial discretization and expanded the range of physics in the simulation codes we use to perform reactor simulations and depletion calculations. This allows us to generate data that are more representative of real environmental samples, especially for the length scale of the isotopic composition and associated variation. Accordingly, these new data allow a more realistic assessment of traditional and new data analytic analysis methods. Here we present motivation for developing reactor simulations using high-performance computing methods and resources, impacts of these new simulations on our assessment of data analysis and interpretation methods, and initial results of developing and systematically testing data analytic methods designed to overcome the challenges expected of real-world samples. We also quantify the performance of these analyses using defensible statistical methods.

Dayman, Ken J.↗

MISPR : an open-source package for high-throughput multiscale molecular simulations

Computational tools provide a unique opportunity to study and design optimal materials by enhancing our ability to comprehend the connections between their atomistic structure and functional properties. However, designing materials with tailored functionalities is complicated due to the necessity to integrate various computational-chemistry software (not necessarily compatible with one another), the heterogeneous nature of the generated data, and the need to explore vast chemical and parameter spaces. The latter is especially important to avoid bias in scattered data points-based models and derive statistical trends only accessible by systematic datasets. Here, we introduce a robust high-throughput multi-scale computational infrastructure coined MISPR (Materials Informatics for Structure–Property Relationships) that seamlessly integrates classical molecular dynamics (MD) simulations with density functional theory (DFT). By enabling high-performance data analytics and coupling between different methods and scales, MISPR addresses critical challenges arising from the needs of automated workflow management and data provenance recording. The major features of MISPR include automated DFT and MD simulations, error handling, derivation of molecular and ensemble properties, and creation of output databases that organize results from individual calculations to enable reproducibility and transparency. In this work, we describe fully automated DFT workflows implemented in MISPR to compute various properties such as nuclear magnetic resonance chemical shift, binding energy, bond dissociation energy, and redox potential with support for multiple methods such as electron transfer and proton-coupled electron transfer reactions. The infrastructure also enables the characterization of large-scale ensemble properties by providing MD workflows that calculate a wide range of structural and dynamical properties in liquid solutions. MISPR employs the methodologies of materials informatics to facilitate understanding and prediction of phenomenological structure–property relationships, which are crucial to designing novel optimal materials for numerous scientific applications and engineering technologies.

36 MATERIALS SCIENCE↗

Elastic Resource Management for Deep Learning Applications in a Container Cluster

The increasing demand for learning from massive datasets is restructuring our economy. Effective learning, however, involves nontrivial computing resources. Most businesses utilize commercial infrastructure providers (e.g., AWS) to host their computing clusters in the cloud, where various jobs compete for available resources. While cloud resource management is a fruitful research field that has made many advances in production, such as Kubernetes and YARN, few efforts have been invested to further optimize the system performance, especially for deep learning (DL) training jobs in a container cluster. This work introduces FlowCon, a system that is able to monitor the individual evaluation functions of DL jobs at runtime, and thus to make placement decisions on resource allocations elastically. Here, we present a detailed design and implementation of FlowCon and conduct intensive experiments over various DL models. The results demonstrate that FlowCon significantly improves DL job completion time and resource utilization efficiency, compared to default systems. According to the results, FlowCon is able to improve the completion time by up to 68.8% and meanwhile, reduce the makespan by 18.0%, in the presence of various DL job workloads.

97 MATHEMATICS AND COMPUTING↗

Performance Characteristics of the BlueField-2 SmartNIC

High-performance computing (HPC) researchers have long envisioned scenarios where application workflows could be improved through the use of programmable processing elements embedded in the network fabric. Recently, vendors have introduced programmable Smart Network Interface Cards (SmartNICs) that enable computations to be offloaded to the edge of the network. There is great interest in both the HPC and high-performance data analytics (HPDA) communities in understanding the roles these devices may play in the data paths of upcoming systems. This paper focuses on characterizing both the networking and computing aspects of NVIDIA’s new BlueField-2 SmartNIC when used in a 100Gb/s Ethernet environment. For the networking evaluation we conducted multiple transfer experiments between processors located at the host, the SmartNIC, and a remote host. These tests illuminate how much effort is required to saturate the network and help estimate the processing headroom available on the SmartNIC during transfers. For the computing evaluation we used the stress-ng benchmark to compare the BlueField-2 to other servers and place realistic bounds on the types of offload operations that are appropriate for the hardware. Our findings from this work indicate that while the BlueField-2 provides a flexible means of processing data at the network’s edge, great care must be taken to not overwhelm the hardware. While the host can easily saturate the network link, the SmartNIC’s embedded processors may not have enough computing resources to sustain more than half the expected bandwidth when using kernel-space packet processing. From a computational perspective, encryption operations, memory operations under contention, and on-card IPC operations on the SmartNIC perform significantly better than the general-purpose servers used for comparisons in our experiments. Therefore, applications that mainly focus on these operations may be good candidates for offloading to the SmartNIC.

97 MATHEMATICS AND COMPUTING↗

Singleton Sieving: Overcoming the Memory/Speed Trade-Off in Exascale k-mer Analysis

Traditional filter data structures, such as Bloom filters, do not offer necessary features that modern high-performance data analytics applications need in order to efficiently perform complex data analysis tasks. For example, MetaHipMer, a de novo metagenome assembler, can use filters to weed out singleton k-mers and reduce memory usage by 30%-70%. However, the filter needs the ability to associate values with k-mers in order to perform the analysis in a single communication pass. Bloom filters do not support value associations and cause the application to perform an extra communication pass, thereby increasing the run time. Therefore, MetaHipMer faces a trade off between memory and speed due to the limited capabilities of traditional filters. In this paper, we overcome the memory and speed trade off in MetaHipMer by integrating a GPU-based feature-rich filter, the Two-Choice filter (TCF), in the MetaHipMer pipeline. The TCF uses key-value association to approximately store k-mers with extensions. This allows MetaHipMer to perform k-mer analysis on the GPUs in a single communication pass. Our empirical analysis shows a 50% reduction in memory usage in k-mer analysis on each node in MetaHipMer without any effect on the overall run time or assembly quality. The memory reduction in turn results in a 43% reduction in the number of nodes required to assemble datasets and enables MetaHipMer to scale to much larger datasets.

McCoy, Hunter↗

Enhanced biochemical sensing with high- Q transmission resonances in free-standing membrane metasurfaces

Optical metasurfaces provide solutions to label-free biochemical sensing by localizing light resonantly beyond the diffraction limit, thereby selectively enhancing light–matter interactions for improved analytical performance. However, high-Q resonances in metasurfaces are usually achieved in the reflection mode, which impedes metasurface integration into compact imaging systems. Here, we demonstrate a metasurface platform for advanced biochemical sensing based on the physics of the bound states in the continuum (BIC) and electromagnetically induced transparency (EIT) modes, which arise when two interfering resonances from a periodic pattern of tilted elliptic holes overlap both spectrally and spatially, creating a narrow transparency window in the mid-infrared spectrum. We experimentally measure these resonant peaks observed in transmission mode (Q ~ 734 at λ ~ 8.8 µm) in free-standing silicon membranes and confirm their tunability through geometric scaling. We also demonstrate the strong coupling of the BIC-EIT modes with a thinly coated PMMA film on the metasurface, characterized by a large Rabi splitting (32 cm -1 ) and biosensing of protein monolayers in transmission mode. Our new photonic platform can facilitate the integration of metasurface biochemical sensors into compact and monolithic optical systems while being compatible with scalable manufacturing, thereby clearing the way for on-site biochemical sensing in everyday applications.

Rosas, Samir [Univ. of Wisconsin, Madison, WI (Uni↗

A Conceptual Framework for HPC Operational Data Analytics

This paper provides a broad framework for under- standing trends in Operational Data Analytics (ODA) for High- Performance Computing (HPC) facilities. The goal of ODA is to allow for the continuous monitoring, archiving, and analysis of near real-time performance data, providing immediately actionable information for multiple operational uses. In this work, we combine two models to provide a comprehensive HPC ODA framework: one is an evolutionary model of analytics capabilities that consists of four types, which are descriptive, diagnostic, predictive and prescriptive, while the other is a four- pillar model for energy-efficient HPC operations that covers facility, system hardware, system software, and applications. This new framework is then overlaid with a description of current development and production deployments of ODA within leading- edge HPC facilities. Finally, we perform a comprehensive survey of ODA works and classify them according to our framework, in order to demonstrate its effectiveness.

Netti, Alessio↗

Visual HPC Workflows for the Analysis of System Dynamics Models

Visual analytics supported by high performance computing (HPC) accelerates and enhances the discovery, exploration, and analysis of causal patterns in complex system dynamics (SD) models. We present a suite of visualization-assisted ensemble-based techniques for hypothesis generation and testing, and for sensitivity analysis. By employing HPC to provide parallel, on-demand simulation of SD models, one can “steer” an ensemble of simulated scenarios in real time as one first formulates and then informally tests those hypotheses: this provides rapid feedback for analysts to refine their understanding of the causal relationships emergent from a model. Such understandings can be followed and augmented by rigorous application of statistical methods, namely global variance-based sensitivity analysis, Monte-Carlo filtering, adaptive regional sensitivity analysis, and self-organized maps: here timely computation relies on HPC, while effective presentation emphasizes high-dimensional multivariate data visualization. Immersive visualization in virtual 3D environments provides an excellent adjunct to the traditional 2D graphics typically used for SD models, as it generates an embodied understanding of model behavior and facilitates an active, collaborative critique of model structure and output. Finally, we summarize prospects for HPC-enabled visual analytics applied to SD modeling.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Tandem Predictions for HPC Jobs

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

HPC↗

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗

pnnl/COMET

COMET: Domain Specific Compilation in Multi-level IR COMET is a compiler for dense and sparse tensor algebra and a domain-specific-language (DSL) that facilitates development and implementation of high-performance computing, graph analytics, and artificial intelligence (AI) applications.

Kestor, Gokcen↗

Analytical and Performance-Based Evaluation Alternatives to Full-Scope, High-Fidelity Testbeds

One consequence the design and operational differences of advanced reactors is that integrated system validation (ISV) using full-scope, high-fidelity testbed might not be cost-justified or practical. The absence of traditional ISV may pose an issue for the conduct of safety evaluations since it may leave regulators without the information derived from this performance-based testing. The purpose of this research is to identify analytical and performance-based test and evaluation methods that are alternatives to full-scope, high-fidelity testbeds that have traditionally been used for ISV. We developed a method to test validation based on a multi-stage validation (MSV) framework. MSV is an approach to meeting validation objectives through incremental, successive validation activities beginning in the early stages of the design process and continuing through the later stages. MSV doesn't rely solely on late stage validation, rather it accommodates the use of a broad spectrum of analytical and performance-based information and a diversity of testbeds. In addition, MSV embraces the use of information from other types of evaluations and analyses.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

FY2020 Energy Efficient Mobility Systems Annual Progress Report

EEMS Program activities during FY 2020 focused on analytical research and large-scale modeling and simulation to understand the impacts that new mobility technologies and services will have at the vehicle-, traveler-, and overall transportation system-level. This research included the development of a multi-fidelity, end-to-end transportation system models and tools to evaluate the complex interactions among the various actors within the mobility landscape, analysis of empirical data to characterize which solutions may provide the largest benefits, and development of new control systems and algorithms that use vehicle connectivity and automation to improve the performance and efficiency of individual vehicles as well as the overall traffic system. This document presents a brief overview of the EEMS Program and documents progress and results from projects within each of the EEMS activity areas. The Computational Modeling and Simulation key activity area summarizes work within the sub-areas of (1) the SMART (Systems and Modeling for Accelerated Research in Transportation) Mobility Lab Consortium, (2) Artificial Intelligence, High-Performance Computing, and Data Analytics, and (3) Core Simulation and Evaluation Tools. Additionally, the program’s advanced R&D projects are summarized within (4) the Connectivity and Automation Technology key activity area. Each of the individual progress reports provide a project overview and highlights of the technical results.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Phoenix: A Scalable Streaming Hypergraph Analysis Framework

We present Phoenix, a scalable hypergraph analytics framework for data analytics and knowledge discovery that was implemented on the leadership class computing platforms at Oak Ridge National Laboratory (ORNL). Our software framework comprises a distributed implementation of a streaming server architecture which acts as a gateway for various hypergraph generators/external sources to connect. Phoenix has the capability to utilize diverse hypergraph generators, including HyGen, a very large-scale hypergraph generator developed by ORNL. Phoenix incorporates specific algorithms for efficient data representation by exploiting hidden structures of the hypergraphs. Our experimental results demonstrate Phoenix’s scalable and stable performance on massively parallel computing platforms. Phoenix’s superior performance is due to the merging of high-performance computing with data analytic.

Kurte, Kuldeep↗

Oak Ridge National Laboratory Annual Sustainability Report 2023

ORNL, managed under contract by UT-Battelle LLC, is DOE’s largest science and energy laboratory and, as such, executes the widest range of mission capabilities. Diverse expertise spans a broad range of scientific and engineering disciplines, enabling research and science achievements to accelerate the delivery of solutions to the marketplace. ORNL supports DOE’s national missions of scientific discovery, clean energy, and security. To execute these activities, ORNL has grown significantly over 80 years of continuous operations, consisting of facilities with commissioning dates ranging from the 1940s to the present—an extraordinary set of distinctive scientific facilities and equipment. The complexities of such a variety of facilities require teamwork among divisions, a wide variety of conservation projects, and creative strategies to achieve the desired energy and water savings. Such a diverse and unique set of major facilities, totaling over 5.5 million square feet, with 6,000 employees, requires an innovative plan to accomplish advancements in operational efficiencies. ORNL is tasked with the management of an extraordinary set of distinctive scientific facilities and equipment for DOE. ORNL is mission-driven, and its mission has grown substantially over the decades. ORNL’s core research capabilities provide broad science and technology support for DOE in the areas of energy, environment, and national security. Currently, ORNL is a world leader in materials, neutron, and nuclear science and engineering, and in high-performance computing and data analytics. ORNL’s vast portfolio of research facilities must be maintained and carefully upgraded to protect the nation’s investment in scientific analysis. The goal of sustainable and resilient operations is to enable more effective execution of ORNL’s science and technology mission. Sustainable operational practices and enhanced resilience strive for excellent results while remaining diligent in energy conservation, environmental stewardship, asset management, and community engagement. The Sustainable ORNL Program (Sustainable ORNL) Continuous improvements in operational and business processes must be integrated into the fabric of the ORNL culture to maximize the return from the investment made in modernizing facilities and equipment. The Sustainable ORNL program promotes the legacy of system-wide best practices, management commitment, and employee engagement that will lead ORNL into a future of efficient, resilient, and sustainable operations. ORNL leadership and Sustainable ORNL champions receive regular status reports on the progress of each project and focus area (i.e., roadmap) and periodic summary reports. More information can be found at the program’s website. The Sustainable ORNL roadmap structure endorses 15 vital roadmaps. The figure below summarizes the current project assignments and demonstrates that each project contributes to the wellbeing of the whole. Continuous employee engagement and regular status reports confirm the ideals of the program. The roadmap structure is not static; as the science mission advances and the needs of the organization evolve, the Sustainable ORNL roadmap structure elements are modified to align with developing priorities. In 2022, Sustainable ORNL made roadmap changes to better align ORNL to support new federal requirements that have been issued.

54 ENVIRONMENTAL SCIENCES↗