Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high performance analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Climatespark: an In-Memory Distributed Computing Framework for Big Climate Data Analytics

The unprecedented growth of climate data creates new opportunities for climate studies, and yet big climate data pose a grand challenge to climatologists to efficiently manage and analyze big data. The complexity of climate data content and analytical algorithms increases the difficulty of implementing algorithms on high performance computing systems. This paper proposes an in-memory, distributed computing framework, ClimateSpark, to facilitate complex big data analytics and time-consuming computational tasks. Chunking data structure improves parallel I/O efficiency, while a spatiotemporal index is built for the chunks to avoid unnecessary data reading and preprocessing. An integrated, multi-dimensional, array-based data model (ClimateRDD) and ETL operations are developed to address big climate data variety by integrating the processing components of the climate data lifecycle. ClimateSpark utilizes Spark SQL and Apache Zeppelin to develop a web portal to facilitate the interaction among climatologists, climate data, analytic operations and computing resources (e.g., using SQL query and Scala/Python notebook). Experimental results show that ClimateSpark conducts different spatiotemporal data queries/analytics with high efficiency and data locality. ClimateSpark is easily adaptable to other big multiple- dimensional, array-based datasets in various geoscience domains.

Hu, Fei↗

Analytical prediction and experimental verification of performance at various operating conditions of a dual-mode traveling wave tube with multistage depressed collectors

A comparison of analytical and experimental results is presented for a high performance dual-mode traveling wave tube (TWT) operated over a wide range conditions. The computations are carried out with advanced multidimensional computer programs. These programs model the electron beam as a series of disks or rings of charge and follow their trajectories from the rf input of the TWT through the slow-wave structure refocusing system to their points of impacts in the depressed collector. TWT performance, collector efficiency, and collector current distribution are computed and compared with measurements. Very good agreement was obtained between computed and measured TWT performance and collector efficiencies, and the computer design of a highly efficient collector was demonstrated.

Dayton, J. A., Jr.↗

Characterization and Valuation of the Uncertainty of Calibrated Parameters in Microsimulation Decision Models

We evaluated the implications of different approaches to characterize the uncertainty of calibrated parameters of microsimulation decision models (DMs) and quantified the value of such uncertainty in decision making. We calibrated the natural history model of CRC to simulated epidemiological data with different degrees of uncertainty and obtained the joint posterior distribution of the parameters using a Bayesian approach. We conducted a probabilistic sensitivity analysis (PSA) on all the model parameters with different characterizations of the uncertainty of the calibrated parameters. We estimated the value of uncertainty of the various characterizations with a value of information analysis. We conducted all analyses using high-performance computing resources running the Extreme-scale Model Exploration with Swift (EMEWS) framework. The posterior distribution had a high correlation among some parameters. The parameters of the Weibull hazard function for the age of onset of adenomas had the highest posterior correlation of -0.958. When comparing full posterior distributions and the maximum-a-posteriori estimate of the calibrated parameters, there is little difference in the spread of the distribution of the CEA outcomes with a similar expected value of perfect information (EVPI) of $\$$653 and $\$$685, respectively, at a willingness-to-pay (WTP) threshold of $\$$66,000 per quality-adjusted life year (QALY). Ignoring correlation on the calibrated parameters’ posterior distribution produced the broadest distribution of CEA outcomes and the highest EVPI of $\$$809 at the same WTP threshold. Different characterizations of the uncertainty of calibrated parameters affect the expected value of eliminating parametric uncertainty on the CEA. Ignoring inherent correlation among calibrated parameters on a PSA overestimates the value of uncertainty.

97 MATHEMATICS AND COMPUTING↗

Spectroscopic Chemical Analysis Methods and Apparatus

This invention relates to non-contact spectroscopic methods and apparatus for performing chemical analysis and the ideal wavelengths and sources needed for this analysis. It employs deep ultraviolet (200- to 300-nm spectral range) electron-beam-pumped wide bandgap semiconductor lasers, incoherent wide bandgap semiconductor lightemitting devices, and hollow cathode metal ion lasers. Three achieved goals for this innovation are to reduce the size (under 20 L), reduce the weight [under 100 lb (.45 kg)], and reduce the power consumption (under 100 W). This method can be used in microscope or macroscope to provide measurement of Raman and/or native fluorescence emission spectra either by point-by-point measurement, or by global imaging of emissions within specific ultraviolet spectral bands. In other embodiments, the method can be used in analytical instruments such as capillary electrophoresis, capillary electro-chromatography, high-performance liquid chromatography, flow cytometry, and related instruments for detection and identification of unknown analytes using a combination of native fluorescence and/or Raman spectroscopic methods. This design provides an electron-beampumped semiconductor radiation-producing method, or source, that can emit at a wavelength (or wavelengths) below 300 nm, e.g. in the deep ultraviolet between about 200 and 300 nm, and more preferably less than 260 nm. In some variations, the method is to produce incoherent radiation, while in other implementations it produces laser radiation. In some variations, this object is achieved by using an AlGaN emission medium, while in other implementations a diamond emission medium may be used. This instrument irradiates a sample with deep UV radiation, and then uses an improved filter for separating wavelengths to be detected. This provides a multi-stage analysis of the sample. To avoid the difficulties related to producing deep UV semiconductor sources, a pumping approach has been developed that uses ballistic electron beam injection directly into the active region of a wide bandgap semiconductor material.

Hug, William F.↗

Symbolic construction of the chemical Jacobian of quasi-steady state (QSS) chemistries for Exascale computing platforms

The Quasi-Steady State Approximation (QSSA) can be an effective tool for reducing the size and stiffness of chemical mechanisms for implementation in computational reacting flow solvers. However, for many applications, the resulting model still requires implicit methods for efficient time integration. Here, in this paper, we outline an approach to formulating the QSSA reduction that is coupled with a strategy to generate C++ source code to evaluate the net species production rates, and the chemical Jacobian. The code-generation component employs a symbolic approach enabling a simple and effective strategy to analytically compute the chemical Jacobian. For computational tractability, the symbolic approach needs to be paired with common subexpression elimination which can negatively affect memory usage. Several solutions are outlined and successfully tested on a 3D multipulse ignition problem, thus allowing portable application across chemical model sizes and GPU capabilities. The implementation of the proposed method is available at https://github.com/AMReX-Combustion/PelePhysics under an open-source license.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Early Exploration of a Flexible Framework for Efficient Quantum Linear Solvers in Power Systems

The rapid integration of renewable energy resources presents formidable challenges in managing power grids. While advanced computing and machine learning techniques offer some solutions for accelerating grid modeling and simulation, there remain complex problems that classical computers cannot effectively address. Quantum computing, a promising technology, has the potential to fundamentally transform how we manage power systems, especially in scenarios with a higher proportion of renewable energy sources. One critical aspect is solving linear systems of equations, crucial for power system applications like power flow analysis, for which the Harrow-Hassidim-Lloyd (HHL) algorithm is a well-known quantum solution. However, HHL quantum circuits often exhibit excessive depth, making them impractical for current Noisy-Intermediate-Scale-Quantum (NISQ) devices. In this paper, we introduce a versatile framework, powered by NWQSim, that bridges the gap between power system applications and quantum linear solvers available in Qiskit. This framework empowers researchers to efficiently explore power system applications using quantum linear solvers. Through innovative gate fusion strategies, reduced circuit depth, and GPU acceleration, our simulator significantly enhances resource efficiency. Power flow case studies have demonstrated up to a eight-fold speedup compared to Qiskit Aer, all while maintaining comparable levels of accuracy.

quantum computing, Harrow-Hassidim-Lloyd, high-per↗

Analytical study of the cruise performance of a class of remotely piloted, microwave-powered, high-altitude airplane platforms

Each cycle of the flight profile consists of climb while the vehicle is tracked and powered by a microwave beam, followed by gliding flight back to a minimum altitude. Parameter variations were used to define the effects of changes in the characteristics of the airplane aerodynamics, the power transmission systems, the propulsion system, and winds. Results show that wind effects limit the reduction of wing loading and increase the lift coefficient, two effective ways to obtain longer range and endurance for each flight cycle. Calculated climb performance showed strong sensitivity to some power and propulsion parameters. A simplified method of computing gliding endurance was developed.

Morris, C. E. K., Jr.↗

Theoretical Combustion Performance of Several High-Energy Fuels for Ramjet Engines

An analytical evaluation of the air and fuel specific-impulse characteristics of magnesium, magnesium octene-1 slurries, aluminum, aluminum octene-1 slurries, boron, boron octene-1 slurries, carbon, hydrogen, alpha-methylnaphthalene, diborane, pentaborane, and octene-1 is presented. While chemical equilibrium was assumed in the combustion process, the expansion was assumed to occur at fixed composition.

Tower, Leonard K↗

Development of the CSI phase-3 evolutionary model testbed

This report documents the development effort for the reconfiguration of the Controls-Structures Integration (CSI) Evolutionary Model (CEM) Phase-2 testbed into the CEM Phase-3 configuration. This step responds to the need to develop and test CSI technologies associated with typical planned earth science and remote sensing platforms. The primary objective of the CEM Phase-3 ground testbed is to simulate the overall on-orbit dynamic behavior of the EOS AM-1 spacecraft. Key elements of the objective include approximating the low-frequency appendage dynamic interaction of EOS AM-1, allowing for the changeout of components, and simulating the free-free on-orbit environment using an advanced suspension system. The fundamentals of appendage dynamic interaction are reviewed. A new version of the multiple scaling method is used to design the testbed to have the full-scale geometry and dynamics of the EOS AM-1 spacecraft, but at one-tenth the weight. The testbed design is discussed, along with the testing of the solar array, high gain antenna, and strut components. Analytical performance comparisons show that the CEM Phase-3 testbed simulates the EOS AM-1 spacecraft with good fidelity for the important parameters of interest.

Gronet, M. J.↗

Experimental and analytical derivation of arc-heater scaling laws for simulating high-enthalpy environments for Aeroassisted Orbital Transfer Vehicle application

The computer code ARCFLO II was used as a guide to increase the performance of the Interaction Heating Facility at Ames Research Center. A closed-form scaling law relation was derived that provides an understanding of the factors that affect enthalpy in the constricted-arc heater. From a study of this scaling law, it is concluded that at constant pressure, enthalpy is proportional to current density raised to the 0.60 power for current densities from 80 to 150 A/sq cm. At constant current density, enthalpy is inversely proportional to pressure to the nth power, where n varies from 0.14 to 0.43, depending on the current density. Radiative heat losses are responsible for the falloff in performance at combinations of high current density and high pressure. An analytical, closed form scaling law based on a constant-temperature arc-core model agrees qualitatively with the scaling law deduced from ARCFLO II.

Winovich, W.↗

Sandia’s Liquid-Cooled Data Center Boosts Efficiency and Resiliency

The Federal Energy Management Program (FEMP) encourages federal agencies and organizations to improve data center energy efficiency, which can offer tremendous opportunities for energy and cost savings. In this success story, a novel liquid cooling system provides reliable, resilient, and energy-efficient cooling for high performance computing (HPC) systems at Sandia National Laboratories.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MOSIQS: Persistent Memory Object Storage With Metadata Indexing and Querying for Scientific Computing

Scientific applications often require high-bandwidth shared storage to perform joint simulations and collaborative data analytics. Shared memory pools provide a chance to satisfy such needs. Recently, a high-speed network such as Gen-Z utilizing persistent memory (PM) offers an opportunity to create a shared memory pool connected to compute nodes. However, there are several challenges to use scientific applications on the shared memory pool directly such as scalability, failure-atomicity, and lack of scientific metadata-based search and query. In this paper, we propose MOSIQS, a persistent memory object storage framework with metadata indexing and querying for scientific computing. We design MOSIQS based on the key idea that memory objects on PM pool can live beyond the application lifetime and can become the sharing currency for applications and scientists. MOSIQS provides an aggregate memory pool atop an array of persistent memory devices to store and access memory objects to accelerate scientific computing. MOSIQS uses a lightweight persistent memory key-value store to manage the metadata of memory objects, which enables memory object sharing. To facilitate metadata search and query over millions of memory objects resident on memory pool, we introduce Group Split and Merge (GSM), a novel persistent index data structure designed primarily for scientific datasets. GSM splits and merges dynamically to minimize the query search space and maintains low query processing time while overcoming the index storage overhead. MOSIQS is implemented on top of PMDK. We evaluate the proposed approach on many-core server with an array of real PM devices. Experimental results show that MOSIQS gains a 100% write performance improvement and executes multi-attribute queries efficiently with 2.7× less index storage overhead offering significant potential to speed up scientific computing applications.

97 MATHEMATICS AND COMPUTING↗

Improving Progressive Retrieval for HPC Scientific Data using Deep Neural Network

As the disparity between compute and I/O on high-performance computing systems has continued to widen, it has become increasingly difficult to perform post-hoc data analytics on full-resolution scientific simulation data due to the high I/O cost. Error-bounded data decomposition and progressive data retrieval framework has recently been developed to address such a challenge by performing data decomposition before storage and reading only part of the decomposed data when necessary. However, the performance of the progressive retrieval framework has been suffering from the over-pessimistic error control theory, such that the achieved maximum error of recomposed data is significantly lower than the required error. Therefore, more data than required is fetched for recomposition, incurring additional I/O overhead. In order to tackle this issue, we propose a DNN-based progressive retrieval framework that can better identify the minimum amount of data to be retrieved. Our contributions are as follows: 1) We provide an in-depth investigation of the recently developed progressive retrieval framework; 2) We propose two designs of prediction models (named D-MGARD and E-MGARD) to estimate the amount of retrieved data size based on error bounds. 3) We evaluate our proposed solutions using scientific datasets generated by real-world simulations from two domains. Evaluation results demonstrate the effectiveness of our solution in accurately predicting the amount of retrieval data size, as well as the advantages of our solution over the traditional approach to reducing the I/O overhead. Based on our evaluation, our solution is shown to read significantly less data (5% - 40% with D-MGARD, 20% - 80% with E-MGARD).

Wang, Jinzhen↗

Hamilton: Flexible, Open Source $10 Wireless Sensor System for Energy Efficient Building Operation

Sensors for improving building performance are rapidly populating the market, driven in part by the drive to reduce greenhouse gas emissions resulting from energy production as well as improve the interior environment for healthy and more productive spaces. UC Berkeley has led wireless sensor development over the past 25 years (e.g., Telos mote), with the Hamilton (named after Alexander Hamilton on the US $10 bill) as the most recent. The Hamilton sensor was designed as a low-cost high-performance sensor that is modular and interoperable. The objective of the Hamilton project was to create, evaluate and establish the technological foundations for secure and easy to deploy building energy efficiency applications utilizing pervasive, low-cost wireless sensors integrated with traditional Building Management Systems (BMS), consumer-sector building components, and powerful data analytics. The project included iterative hardware design, incorporating a high-performance database (BTrDb, http://btrdb.io/), creating and iterating the development of secure data middleware (BOSSwave, WAVE/WAVEMQ), working with and pushing the development of an open-source tiny operating system RiotOS, and implementing and improving protocols such as Thread/OpenThread and TCP/IP. The hardware benefited from careful design to drive down the cost; the design included a System-on-a-Chip (SoC), chip antenna, single crystal and five passive components. Careful design of the operating system created a low-power design to enable a long life with small batteries. The hardware included several sensors: temperature, radiant temperature, relative humidity, magnetometer, accelerometer, and light, with an optional occupancy (Passive InfraRed) sensor. The project was the basis of several applications, both internal to the research team and other researchers and professionals at other institutions. Several applications used the sensor hardware as the basis for other complex devices. Other applications used the sensors to improve building performance through interoperating with the building Heating Ventilation and Air-Conditioning (HVAC) system, such as using occupancy and/or distributed temperature sensing to reduce HVAC zone energy while still providing thermal comfort and to reduce peak loads in small commercial buildings. We demonstrated cloud-based energy analytics, implemented a schedule and a Model Predictive Controller in a small commercial building to optimize HVAC energy, occupancy and electricity price. Initial integration of these technological innovations was performed through the creation of execution containers containing the WAVE agent and various driver, proxy, or building system function logic. The research added to the understanding of efficient sensor hardware, secure middleware, time-series data management (high performance database), efficient communication protocols, and interoperating with applications and building systems. The project showed the technical effectiveness and economic feasibility of creating a low-cost, modular, and easy-to-deploy sensor. Through conversations with multiple end users, the research team discovered that many customers wanted data management and services in addition to the sensors. HamiltonIOT developed packages of sensors, border router, and data services to provide a seamless “plug-and-play” sensor deployment. Some customers were willing to pay for higher quality sensors (such as light); some customers wanted a robust enclosure (waterproof).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Calculated performance map of a 4 1/2-stage 15.0 centimeter (5.9 inch) mean diameter turbine designed for a turbofan simulator

The overall performance of an existing high-ratio turbine is calculated analytically over a range of speed and pressure ratio in order to determine its capability for other applications. The analytical performance covers a speed range from 50 to 120 percent of design and a pressure-ratio range from 5.0 to 35.0. The turbine was designed for a 50.8 centimeter (20.0 in.) tip diameter turbofan simulator. Computed results are compared with the experimental turbine data obtained from testing three fan configurations with the turbofan simulator in air. The comparison indicates good agreement over the range of speeds and pressure ratios covered by the experimental data.

Wasserbauer, C. A.↗

STREAM: A Scalable Federated HPC Telemetry Platform

Obtaining and analyzing high performance computing (HPC) telemetry in real time is a complex task that can impact algo- rithmic performance, operating costs, and ultimately scientific outcomes. If your organization operates multiple HPC systems, filesystems, and clusters, telemetry streams can be synthesized in order to ease operational and analytics burden. In order to collect this telemetry, the Oak Ridge Leadership Computing Facility (OLCF) has deployed STREAM (Streaming Telemetry for Resource Events, Analytics, and Monitoring), which is a distributed and high-performance message bus based on Apache Kafka. STREAM collects center-wide performance information and must interface with many sources, including five HPE deployed supercomputers, each with their own Kafka cluster which is managed by HPCM. OLCF Supercomputers and their attached scratch filesystems currently send more than 300 million messages to over 200 topics producing around 1.3 Terabytes per day of telemetry data to STREAM. This paper describes the architectural principles that enable STREAM to be both resilient and highly performant while supporting multiple upstream Kafka clusters and other data sources. It also discusses the design challenges and decisions faced in adapting our existing system- monitoring infrastructure to support the first Exascale computing platform.

Adamson, Ryan↗

Investigation of arterial gas occlusions

The effect of noncondensable gases on high-performance arterial heat pipes was investigated both analytically and experimentally. Models have been generated which characterize the dissolution of gases in condensate, and the diffusional loss of dissolved gases from condensate in arterial flow. These processes, and others, were used to postulate stability criteria for arterial heat pipes under isothermal and non-isothermal condensate flow conditions. A rigorous second-order gas-loaded heat pipe model, incorporating axial conduction and one-dimensional vapor transport, was produced and used for thermal and gas studies. A Freon-22 (CHCIF2) heat pipe was used with helium and xenon to validate modeling. With helium, experimental data compared well with theory. Unusual gas-control effects with xenon were attributed to high solubility.

Saaski, E. W.↗